To use AI in business, work in this order: list the AI already doing work in your company, give that work one approved route to a model, promote one item on the list into an owned asset with a named owner, put a gate on the single step you cannot undo, then attach a monthly cost figure to it. Five stages, five artifacts, one at a time.
Our position is that the sequence matters more than the use case, and that almost every guide gets the first stage wrong. The standard AI strategy advice is to run a workshop, score candidate use cases on impact and feasibility, and pilot the winner. We think that is the second thing you do, not the first. In most companies the useful AI workflow already exists, unofficially, on somebody's personal account, and the workshop is how you fail to find it.
A network engineer on Hacker News described the shape of this exactly. Writing about joining a power company where he had spotted a clear AI opportunity in drone-based equipment inspection, taha_moji said "the environment made it hard to move fast". The systems were outdated and there was no support for building AI in-house. He had the use case. He did not have a route, an owner, or an approval path, and the idea stalled there. That is the ordinary failure, and it is a sequencing failure, not an imagination failure.
The short answer: Sequence AI adoption by artifact — use list, route record, scope sheet, approval log, cost line — and refuse to start any stage until the previous stage's artifact exists and someone can produce it in ten minutes.
Last updated: July 30, 2026.
The sequence, with the artifact that gates each step. The artifacts are the deliverables; the stages are just labels for the work between them.
One disclosure before the method: we have not run this exact five-stage sequence end to end inside a customer and measured the result, so nothing below is presented as our own before-and-after number. What follows is a procedure built from published research, the failure patterns we see in AI deployment work, and the control surfaces our own platform exposes.
What "Using AI in Business" Actually Means in 2026
Using AI in business now means one of three distinct things, and the confusion between them wastes more time than any technical problem. It can mean employees using a chat assistant to draft and summarise, which nearly everyone already does. It can mean a defined workflow where a model reads input, decides something, and writes a result into a system of record. Or it can mean AI agents that plan and execute several steps against live systems. Only the second and third change how the business runs.
The gap between those meanings shows up in the survey data. McKinsey's global survey published 5 November 2025 found 88 percent of respondents report regular AI use in at least one business function, up from 78 percent a year earlier — but only about one-third say their organisation has begun to scale AI, and just 39 percent report any EBIT impact at the enterprise level, with most of those putting it below 5 percent of EBIT. Stanford HAI's 2026 AI Index puts organisational adoption at 88 percent for 2025 from an independent sample, which is a rare case of two large studies landing on the same figure.
So adoption is close to universal and impact is not. That is the actual problem you are solving when you ask how to use AI in business. You are not trying to be one of the 88 percent — you already are. You are trying to be one of the third that got past pilots. The difference between those groups is structural rather than technological, which is the whole reason sequencing is the subject of this guide rather than tool selection.
| What people mean by "using AI" | Who does it | Changes the P&L? |
|---|---|---|
| Assistants for drafting, summarising, research | Individuals, ad hoc | Rarely visible; the time saved is reabsorbed |
| A defined workflow: read, decide, write to a system | A team, on a schedule or a trigger | Yes, if the write step is trusted |
| An agent taking multi-step action across systems | A team, with an owner and limits | Yes, and it is where the risk concentrates |
McKinsey's own read on the differentiator is worth quoting precisely, because it is not about models. The organisations seeing real value are nearly three times as likely to say they have fundamentally redesigned individual workflows, and that redesign had one of the strongest measured contributions of all factors tested. They are also more likely to have defined processes for deciding when a model's output needs human validation. Redesigned workflows and defined validation points. That is a description of artifacts, not of a technology choice.
Why Use-Case-First Sequencing Fails
Use-case-first sequencing fails because the selection step runs on opinion while the useful data already exists in your own building. When a company opens with a prioritisation workshop, it produces a ranked list of things nobody has tried, sorted by guesses about impact. Meanwhile the tasks people have actually automated with AI, successfully, sit unrecorded on personal accounts.
The scale of that unrecorded activity is the most under-used fact in the field. MIT's Project NANDA report, The GenAI Divide: State of AI in Business 2025, found that while only 40 percent of companies said they had purchased an official LLM subscription, workers from over 90 percent of the companies surveyed reported regular use of personal AI tools for work tasks. The report's own section on this is blunt: employees are already crossing the divide through personal tools, and that shadow activity "often delivers better ROI than formal initiatives and reveals what actually works."
Independent telemetry points the same way. Harmonic Security's May 2026 analysis of 1.9 million classified AI-session minutes found that 64.5 percent of activity on personal and free-tier accounts is business use, not personal use — 60.6 percent on ChatGPT Free, 67.2 percent on Claude Free, 80.2 percent on Copilot Free. This is vendor telemetry from one security tool's install base rather than a random sample, and should be read as a strong directional signal rather than a population estimate. It still describes something a workshop cannot see.
The consequence is that use-case-first selection is not merely slower. It is systematically biased. The MIT report found roughly half of generative AI budget allocated to sales and marketing, while back-office automation often yields better return, and attributes the bias to easier metric attribution rather than actual value. Customer support and content generation get funded; invoice and dispatch integration work does not. Companies pick what is easy to claim credit for. An inventory of real usage does not have that bias, because it records what people chose to do when nobody was measuring them.
There is a second reason to start with observation, and it is the one that decides whether stage three ever happens. A workflow discovered on a personal account has already been validated by the only test that matters: somebody used it repeatedly, voluntarily, to get their own work done. A workshop candidate has passed no test at all. That is the difference between an AI use case you found and one you invented. Our earlier analysis of why AI pilots stall on the way to production sets out the four questions a pilot never has to answer and production always does; a promoted workflow arrives at that gate with two of them already answered, because it has a de facto owner and a demonstrated volume.
A practitioner on Hacker News reading the same MIT report reached the conclusion directly, noting that "almost everyone was using AI tools they paid for" themselves, and adding the sharper instruction: read the reports, not just headlines. The finding was available to anyone who read past page one.
Prerequisites: Four Things That Must Be True Before Stage One
Before stage one, four conditions need to hold, and none of them is technical. If any is missing, the sequence stalls in a predictable place, so it is cheaper to check them in an afternoon than to discover them in month three.
A named person who owns the sequence itself. Not a committee and not "IT". One person whose job includes producing the five artifacts and who can be asked, in a corridor, what stage you are on. Where this is missing, stage one produces a list nobody updates.
Permission to ask without punishment. Stage one asks employees to disclose that they have been pasting work into tools you never approved. If disclosure risks a disciplinary conversation, you will get a short and flattering list. State in writing, before you ask, that nobody is in trouble for what they report this week and that the point is to make the tool legitimate rather than to remove it. This is the single condition most often skipped and it invalidates the whole first artifact.
Somebody who can say yes to a connection. By stage two you need a decision about credentials and network access to a model endpoint. If the only path to that decision is a quarterly architecture board, note the date now and plan the sequence around it.
One system of record you are willing to let AI write into, eventually. Not now — at stage four. But if the honest answer is "none, ever", the sequence will terminate at stage three with a drafting assistant, which is a legitimate outcome and worth knowing in week one rather than week twelve.
Nothing on that list requires a budget line, a data lake, or a strategy document. Two of the four are permissions, and both are usually granted faster than people expect once someone actually asks.
The Five-Artifact Sequence, In Order
There are five stages in the sequence to use AI in business, and each one exists to produce exactly one artifact. The rule that makes it work is the advance rule: you may not start stage N+1 until stage N's artifact exists and someone can produce it in ten minutes. An artifact nobody can find in ten minutes does not exist, whatever the project plan says.
| Stage | The question it answers | Artifact | Typical elapsed time |
|---|---|---|---|
| 1. Observe | Which tasks are already done with AI, by whom, on which account? | The use list | 1-2 weeks |
| 2. One route | Where does a request go, and who is allowed through that door? | The route record | 2-4 weeks |
| 3. Promote one | Who owns this, what may it touch, when is it reviewed? | The scope sheet | 1-2 weeks |
| 4. Commit gate | Which single action is irreversible, and who approves it? | The approval log | 1 week |
| 5. Cost line | What did this cost last month, out of whose budget? | The cost line | ongoing |
Three properties of this design are worth naming, because they are what separate it from a maturity model.
It is artifact-gated rather than time-gated. Nothing advances because a quarter ended. Each gate is a document someone can hand you, which means progress is falsifiable — a claim that you are "in stage four" is checkable in ten minutes by asking for the approval log.
It scales by repetition, not by expansion. Having completed the cycle once for one workflow, you run it again for the second, and the second run reuses stages two and five entirely. The first workflow costs eight to twelve weeks. The fifth costs days. This is why resisting the urge to launch five workflows at once is not caution, it is arithmetic.
It puts the irreversible step last, deliberately. Stages one to three are all reversible: a list can be deleted, a route can be closed, a scope can be narrowed. Stage four is the first point where the AI does something to the world. Reaching it fourth rather than first is the entire safety argument. It costs nothing in speed, because the earlier stages were never the slow part.
The shape will look familiar if you have read our gateway material. LeapForce's rollout guide for the AI Gateway states the same ordering principle in three words, "Observe first. Enforce second. Optimize third.", across four phases: visibility, protection, enforcement, optimisation. The sequence in this article is the business-process version of that idea, generalised past the gateway to the workflows and owners around it.
Stage 1: The Use List
The use list is a table of every task in your company that is currently being done with AI, who does it, which tool, on which kind of account, and what data goes in. It takes one to two weeks and it is produced by asking people, not by scanning traffic. Fifteen conversations of ten minutes each will beat any discovery tool you can procure in the same period.
Ask each team lead one question: what have you or anyone on your team used an AI tool for in the last month, for work? You are looking for the business tasks they already automate, however partially. Then ask the follow-up that matters: was that on a company account or your own? Record the answers in five columns.
| Column | Why it is there |
|---|---|
| Task | Named at the level of a job, not a capability. "Extract lane and weight from inbound quote emails", not "text extraction". |
| Person | The de facto owner. This is your candidate owner at stage three. |
| Tool and account type | Company, paid personal, or free personal. This column is the whole risk picture. |
| Data that goes in | Customer names, contract text, financials, code, HR records. Be specific. |
| Frequency | Times per day or week. Your volume estimate for stage five comes from here. |
Two things reliably happen. First, the list is longer than leadership expects — the MIT and Harmonic figures above predict that, and every discovery exercise we are aware of confirms the direction. Second, the highest-frequency item is almost never the item that would have won a prioritisation workshop. It is usually unglamorous, back-office, and owned by someone junior.
The data column is where the list stops being an inventory and becomes a risk register. A July 2026 survey of 500 employed US adults commissioned by Kolmogorov Law and run through Pollfish, reported by Stacker, found 38 percent had entered at least one type of work information into a personal AI account their employer does not control: 23 percent internal emails or documents, 11.8 percent customer information, 11.4 percent contracts, 10.6 percent HR records. Margin of error was roughly plus or minus 4.4 points. The same survey found 40 percent say their employer has no written AI policy at all.
You are not going to fix that with a policy memo, and the memo is not what stage one is for. The list's job is to convert an unknown into a named set of items you can move, one at a time, onto something you control. We wrote about the broader pattern in shadow AI and why half a company ends up on ungoverned tools; the operational point here is narrower. Do not treat the list as evidence of misconduct. Treat it as a free requirements document that your staff wrote for you.
The artifact is done when: one person can open a single table with at least ten rows, every row has all five columns filled, and the rows are sorted by frequency.
Stage 2: The Route Record
The route record names one approved way for AI work to reach a model, and states who may use it. It exists so that stage three has somewhere to put the workflow you promote. Without it, promoting a workflow means blessing somebody's personal subscription, which is not governance, it is paperwork.
A route is three decisions, and they can be made in a single meeting once someone with authority is in the room.
- Where requests go. A single endpoint, whether that is a vendor's enterprise tier, a gateway you run, or a cloud provider's model service. The requirement is not sophistication. It is singularity: one address that every approved AI call passes through, whatever no-code builder or SDK sits in front of it.
- Who may use it. Tied to your existing identity system, so that the answer to "who has AI access" is a group membership rather than a spreadsheet, and removing someone from the group removes their access everywhere at once.
- What the route records. At minimum: who called, which model, when, and how much it cost. If the route cannot answer those four questions, stages four and five have no data to work with, and you will be rebuilding this in six months.
The order here is a deliberate inversion of common practice. Most companies buy licences first and think about the route later, which is why a mid-sized company can end up with three separate AI billing relationships, none of which knows about the others. Establishing the route before the second workflow means every subsequent one inherits it for free.
There is an obvious objection: does insisting on one route slow everyone down while it is being built? It can. The mitigation is to keep the route deliberately unrestrictive at first. Our own rollout guidance is that phase one is observation with no rules — point traffic at the route, watch what happens, and add nothing. The gateway page describes that first phase as producing truth rather than control: within days you know which tools, models and prompts are in use and what they cost. Enforcement is phase three for a reason. A route that blocks things in week one gets routed around in week two.
The artifact is done when: one person can state the endpoint, name the identity group that may use it, and show a log line for a real call made through it yesterday.
Stage 3: The Scope Sheet
The scope sheet turns one item from the use list into a company asset: a named owner, an explicit list of what the workflow may read and write, and a review date. It is the shortest stage and the one that changes the legal and operational character of what you are doing, because after it the workflow belongs to the company rather than to a person.
Pick the item from stage one with the highest frequency and the narrowest data footprint. High frequency because it will produce enough runs to learn from within a month. Narrow data because the first promotion should not also be your first data-classification argument.
Then write down five fields. Not a document — five fields.
| Field | Example content | Why it is load-bearing |
|---|---|---|
| Owner | A named person, plus their manager as deputy | Somebody has to answer for output quality on the Tuesday after launch |
| Purpose | One sentence, in the business's own words | Scope creep is detected by comparing behaviour to this sentence |
| May read | Named systems and named subsets: "the quotes@ mailbox, unread items only" | Least privilege is only real when it is enumerated |
| May write | Named systems, named fields, and explicitly what it may not do | The gap between read and write is where stage four goes |
| Review date | A date within 90 days | An unreviewed scope silently becomes a permanent one |
The "may write" field deserves the most argument. Restricting a workflow to reads makes it safe and often useless. The value in most business automation is precisely in the write. The resolution is not to choose between them but to split the write into propose and commit: the workflow may write a draft, a proposed value, a queued record, and something else performs the commit. That is stage four's subject, and designing for it here saves rework.
Ownership is the field that survives everything else. Our earlier analysis of moving personal prompts into shared, owned AI agents makes the case that an agent without an owner is not an asset, and that the ownership question becomes urgent at exactly the moment nobody is thinking about it: when the person who built it leaves. A scope sheet with a real name in the owner field is what makes a leaver's departure an administrative event rather than an outage. The related question of what an agent's own identity should look like — owner, scope, expiry, one-step revocation — is covered in our piece on non-human identity for AI agents.
The artifact is done when: one page exists with all five fields filled, the named owner has read it and agreed in writing, and the review date is in a calendar.
Stage 4: The Commit Gate
The commit gate is a control on the one action in the workflow that cannot be undone, plus a log of what was proposed, what was approved, and what was refused. Everything upstream of that action can stay fast and unsupervised. This is the stage that lets you be relaxed about the rest.
Start by finding the irreversible step, which is usually easier than teams expect because there is normally exactly one. Ask: if this ran wrong at three in the morning and nobody noticed until nine, what would we be unable to take back? An email that left the building. A refund issued. A candidate rejected. A record overwritten with no prior version. Those are the ones that need a gate. Reading, drafting, scoring, summarising and queueing are all reversible and none of them needs a gate.
Then choose the gate type honestly against volume.
| Gate type | Use when | Cost |
|---|---|---|
| Human approval on every commit | Volume under roughly 50 per day, or the error is expensive or public | One person's attention; becomes rubber-stamping above that volume |
| Rule-based auto-approve with human exceptions | Most cases fit a stated numeric or categorical bound | Writing the bound honestly; the exception queue must be someone's job |
| Sampled review after the fact | The action is cheap to reverse and volume is high | You will find errors late; unacceptable for anything affecting a person's rights |
| No gate | Genuinely reversible actions only | Zero, and this should be most of your steps |
The reason to be strict here is that the failure mode is well described. OWASP's Top 10 for LLM Applications 2025 names Excessive Agency as LLM06 and traces it to three root causes: excessive functionality, excessive permissions, and excessive autonomy. Its listed example of excessive autonomy is precisely a system that "fails to independently verify and approve high-impact actions". Its prevention list leads with minimising the extensions and functions an agent can call, minimising the permissions those extensions hold on downstream systems, and executing in the user's own security context rather than under a shared high-privilege identity.
There is a design point in that list worth stating separately: the gate belongs downstream of the model, not inside the prompt. OWASP's guidance is to implement authorisation in downstream systems rather than relying on the model to decide whether an action is allowed. An instruction in a system prompt saying "always ask before sending" is not a control, because the thing being asked to obey it is the thing that might malfunction.
The log matters as much as the gate. Record what the workflow proposed, what a human or rule decided, who decided it, and when — including the refusals. Refusals are the more useful half of the record. They are the evidence that the control was doing something, and they are what an auditor or a regulator actually asks about. We have written separately on audit trails that prove what an agent did and on when approval genuinely constitutes control rather than theatre.
The artifact is done when: you can show ten consecutive commit decisions, with at least one refusal among them, and say who made each one.
Stage 5: The Cost Line
The cost line is a single monthly figure for what this workflow costs, attached to a named owner's budget. It is the last artifact because it is the one that governs whether you get to build the next nine, and because it needs a month of real runs before it means anything.
Deriving it is arithmetic, and the arithmetic is deliberately simple.
- Take the frequency from your stage-one row: runs per day.
- Multiply by working days per month.
- Multiply by the cost per run that your stage-two route recorded.
- Add the human time in the stage-four gate, at a loaded hourly rate.
- Attribute the total to the stage-three owner's cost centre.
Step three is the one people skip. It is why AI spend surprises finance. Read the cost per run off your own route's logs rather than estimating it from a price list, because real prompts carry context, retries and failures that a per-token price does not predict. Step four is the one people forget entirely, and on a human-approval gate it is usually the larger number of the two.
What the figure is for is not cost control in the first month — the amounts are typically trivial. It is for the conversation at workflow number ten, when someone in finance asks what all of this costs. The honest answer needs to be a number per owner, not one invoice from a model provider. Budgets expressed per team and per agent, in currency rather than tokens, are what make that answer possible; our write-up on how model routing cuts AI costs covers the mechanism for reducing the figure once you can see it. Note the ordering constraint though: you cannot route on cost before you can measure cost, which is exactly why optimisation is third in our rollout guide and not first.
The artifact is done when: a named owner can state last month's figure for the workflow, and that figure appears somewhere a finance person looks.
A Worked Example: One Freight Company's First Ninety Days
This is a constructed illustration, not a LeapForce customer and not a case study. The company, volumes and figures are invented to show the five artifacts assembled; only the method and the arithmetic are real. We say that plainly because a fabricated customer story would be worse than no example at all.
Northwind Freight, 140 staff, three depots, a transport management system and a shared quotes mailbox. The operations director is given the sequence and one instruction: produce the artifacts.
Weeks 1-2, the use list. Eleven ten-minute conversations produce nine rows. Sorted by frequency, the top four:
| Task | Person | Tool / account | Data in | Frequency |
|---|---|---|---|---|
| Extract lane, weight and date from inbound quote emails | Dispatch supervisor | ChatGPT Free, personal | Customer names, addresses, cargo detail | ~25/day |
| Rewrite carrier claim letters | Claims clerk | ChatGPT Plus, personal paid | Claim text, incident detail | ~6/week |
| Summarise driver incident reports | Depot manager, two of three | Copilot Free, personal | Employee names, incident detail | ~4/week |
| Draft job adverts and screen CVs | HR coordinator | Gemini Free, personal | Applicant CVs | ~10/month |
Leadership had expected the list to be about marketing copy and customer support. It is about dispatch. Two rows carry personal data the company did not know had left its systems, and the CV row is flagged immediately as the one to leave alone until the legal position is settled.
Weeks 3-6, the route record. IT stands up one enterprise endpoint tied to the existing identity provider, with an access group called ai-users containing 31 people. Logging captures caller, model, timestamp and cost per call. No restrictions are applied in this period, on purpose. A two-week delay occurs waiting for a mailbox permission that nobody had realised sat with a different team — the single largest cause of slippage in the whole ninety days, and the reason the prerequisites section above exists.
Weeks 7-8, the scope sheet. The dispatch row is promoted.
| Field | Content |
|---|---|
| Owner | Dispatch supervisor; deputy is the operations director |
| Purpose | Turn an inbound quote email into a structured draft quote request in the TMS, so a human quotes faster |
| May read | The quotes@ mailbox, unread items from the last 7 days, message body and attachments |
| May write | One TMS record in status draft-unpriced, fields: lane, weight, dimensions, requested date, source message id. May not price. May not send. May not email the customer. May not touch existing records. |
| Review date | 90 days from go-live |
Week 9, the commit gate. The irreversible act is the quote leaving the building with a price on it, so that is where the gate goes: the workflow produces draft-unpriced records only, and a human prices and sends. Gate type is human approval on every commit — at 25 a day that is well inside the range where per-item review stays real rather than becoming a reflex. The log records the proposed extraction, the human's corrections, the accept or reject, and the identity of whoever decided.
In the first fortnight the log shows 247 proposals, 231 accepted with no change, 14 corrected, and 2 rejected outright because the inbound email was a duplicate thread the extraction had merged. Those two rejections are the most valuable rows in the log, because they define a failure mode nobody predicted and they prove the gate is load-bearing.
Weeks 10-13, the cost line. The arithmetic, with an explicitly hypothetical unit cost that you must replace with your own measured figure:
| Line | Calculation | Monthly |
|---|---|---|
| Model calls | 25 runs/day x 22 working days x $0.011 per run (read from the route log) | $6.05 |
| Gate time | 550 reviews x 40 seconds = 6.1 hours, at $34/hour loaded | $207.40 |
| Total, attributed to dispatch | $213.45 |
The interesting number is not the total. It is the ratio. Human review is 97 percent of the cost of this workflow. That single figure reframes the next decision correctly. The way to make this cheaper is not a smaller model — it is moving the 231-of-247 uncontested cases onto a rule-based auto-approve with an exception queue, which is a stage-four change, not a stage-two one. Without the cost line, the same team would have spent the next quarter shopping for cheaper inference and saved four dollars.
At day 90 Northwind has all five artifacts, one live workflow, and eight remaining rows on the use list that now cost days each rather than months, because stages two and five are already built.
How Fast The Sequence Can Actually Run
The first cycle takes eight to twelve weeks in a company with normal permissions friction, and the constraint is almost never model quality or engineering effort. It is two waits. One is for someone authorised to approve a connection. The other is for a month of runs to accumulate, so that the cost line means anything.
| Company type | First cycle | What dominates | Where it stalls |
|---|---|---|---|
| Under 20 people, no IT function | 3-5 weeks | Nothing; the owner is also the approver | Stage 5, because volumes are too low to measure for a month |
| 20-250 people, one IT generalist | 8-12 weeks | Waiting on credentials and mailbox or CRM permissions | Stage 2 |
| 250+ with security review and change control | 12-20 weeks | Security review of the route; data classification of the workflow | Stage 2, then stage 3's "may read" line |
| Regulated sector, personal data in scope | 20+ weeks | Legal review, DPIA, records of processing | Stage 3, and correctly so |
Two levers genuinely compress this. Neither is a tool purchase. Running stages one and two in parallel is safe, because the use list does not depend on the route existing — start the conversations in week one and the credential request the same day. And choosing a first workflow whose data footprint is entirely internal, non-personal and low-sensitivity removes the legal review from the critical path of the first cycle entirely; you can take on the harder data in cycle two, when the machinery already exists.
One thing that does not compress it: skipping a stage. The MIT data offers a caution here worth taking seriously — the report found that internal builds fail roughly twice as often as external ones, which is consistent with teams that jump straight from an idea to a build without the intermediate artifacts. McKinsey's high-performer profile points the same direction: the differentiators were redesigned workflows and defined human-validation processes, which are precisely stages three and four.
Eight Ways Companies Break The Sequence
Most failures are not exotic. They are one of eight substitutions, each replacing an artifact with something that resembles it.
1. Starting with a tool selection. Choosing a platform before stage one means the use list gets filtered to what the platform does. The list should be written with no tool in mind, because its job is to tell you what you need.
2. Running the discovery as an audit. If stage one is framed as a compliance sweep, the returns are incomplete and defensive. The framing that works is: tell us what is helping you, so we can make it official and stop you carrying the risk personally.
3. Treating a policy document as stage two. A written AI policy is useful and it is not a route. The test is mechanical: can you name one endpoint and show a log line? A policy produces neither.
4. Promoting three workflows at once. The second and third promotions cost almost nothing once the first is done, and almost everything if done simultaneously, because the first promotion is where you discover what your scope sheet template is missing.
5. Putting the gate at the wrong step. Requiring approval for reads and drafts adds friction with no risk reduction, and it trains people to click through, which is what destroys the gate on the step that mattered. Gate the irreversible act and only that.
6. Building the gate as a prompt instruction. Covered above, and worth repeating because it is common: an instruction to the model is not an authorisation control. OWASP's guidance is explicit that authorisation belongs in downstream systems.
7. Leaving the owner field as a team name. "Owned by Operations" means owned by nobody on the day the output is wrong. A team can hold a budget. Only a person can hold an answer.
8. Skipping the cost line because the amount is small. The first workflow's model spend genuinely is trivial. The artifact is not there to control that month's spend; it is there so that the tenth workflow arrives in a world where per-owner attribution already exists. Retrofitting attribution across ten workflows and three billing relationships is a project. Establishing it on the first one is a spreadsheet row.
A ninth failure sits slightly apart, because it is a leadership pattern rather than a step error. Companies that set out to use AI in business with efficiency as the sole objective tend to pick badly. McKinsey found 80 percent of respondents set efficiency as an objective of their AI work, while the organisations seeing the most value more often set growth or innovation as additional objectives. An efficiency-only frame tends to select for the tasks that are easiest to count, which is the same bias that sends budget to sales and marketing.
What Regulation Asks Of You Right Now
For most companies deploying AI in ordinary business processes, the near-term regulatory ask is smaller than the headlines suggest and it lands mainly on two things you would produce anyway: staff who understand the tools, and records of what the system did. The dates moved in 2026, so it is worth being precise.
Under the EU AI Act, Article 4 has applied since 2 February 2025 and requires providers and deployers to "take measures to ensure, to their best extent, a sufficient level of AI literacy of their staff and other persons dealing with the operation and use of AI systems on their behalf". Note the word deployers: this obligation reaches companies that merely use AI systems, not only those building them.
The high-risk regime moved. Under the Digital Omnibus agreement reached on 7 May 2026, as reported by Pinsent Masons, obligations for stand-alone high-risk AI systems shift from 2 August 2026 to 2 December 2027, and for high-risk AI embedded in regulated products to 2 August 2028. Generative AI transparency requirements were not deferred and remain on the 2 August 2026 date, with a grace period to 2 December 2026 for labelling and watermarking on models placed on the market before then.
| Date | What applies | Relevance to this sequence |
|---|---|---|
| 2 Feb 2025 (in force) | Prohibitions; Article 4 AI literacy for providers and deployers | Stage 1 conversations are a plausible start on literacy; record that you had them |
| 2 Aug 2026 | Generative AI transparency requirements, not deferred | Stage 4's log is where disclosure and traceability live |
| 2 Dec 2026 | Labelling and watermarking grace period ends for earlier models | Relevant if you generate customer-facing content |
| 2 Dec 2027 | Stand-alone high-risk systems | If your stage-3 workflow touches recruitment, credit or similar, this is your clock |
| 2 Aug 2028 | High-risk AI embedded in regulated products | Product teams, not process owners |
Voluntary frameworks line up with the sequence more neatly than they usually do. NIST's AI Risk Management Framework, version 1.0, released 26 January 2023, is organised around four functions: Govern, Map, Measure and Manage. Stage one is Map, stages two to four are Govern and Manage, and stage five is Measure. ISO/IEC 42001 is the other name that comes up in these conversations; we describe it by name only here because the standard's text sits behind a paywall and iso.org blocked every fetch we attempted, so we are not attributing any specific requirement to it. Our longer treatment of the deployer's position is in our EU AI Act guide for AI deployers.
The practical reading: none of this requires a compliance programme before stage one. It requires that by the time an AI system is doing something consequential, you can say who owns it, what it may touch, what it did, and that the people operating it were trained. Which is the five artifacts, described in a different vocabulary.
What This Sequence Will Not Tell You
The sequence is a sequencing tool. It is deliberately silent on several things a reader might reasonably want from an article about using AI in business, and it is more useful if we say which.
It will not tell you whether a task is a good candidate on quality grounds. The use list tells you what people already do successfully, which is a strong signal and not a complete one. Some frequent uses are frequent because they are easy, not because they are valuable. Frequency is a productivity signal, not a value ranking. Our companion piece on why AI business automation starts with rules rather than models covers the separate question of where in a process the model actually belongs.
It will not produce a return-on-investment case in the first cycle. One workflow's cost line is a cost, not a return. Time saved in the stage-four gate is real, but attributing it to the P&L honestly takes several cycles, and the McKinsey figure of 39 percent reporting any EBIT impact, most of them below 5 percent, is the realistic base rate to plan against.
The evidence base for the shadow-AI premise is directional, not definitive. We lean on it, so its weaknesses should be visible. The MIT NANDA report is the most-cited source for the claim and it is also the one we trust least on precision: its own text gives two different figures for the same measure, stating in one place that half of generative AI budgets go to sales and marketing and in another that those functions captured approximately 70 percent of allocation in the same survey. We cite its shadow-AI findings because they are specific and internally consistent, and we deliberately do not build on its headline "95 percent get zero return" figure, which has been widely contested and rests on a small interview and survey base. Harmonic's telemetry is a vendor's own install base. The Kolmogorov Law survey is 500 respondents with a plus-or-minus 4.4-point margin. Three independent methods pointing the same way is a strong argument about direction. It is a weak one about magnitude.
We have not measured this sequence ourselves. Stated once at the top and repeated here because it matters: there is no LeapForce before-and-after study behind the eight-to-twelve-week figure. It is an estimate from observed permission-friction patterns, and the honest way to test it is to run it and time your own stages.
It says nothing about the hardest cases. Anything touching recruitment decisions, credit, healthcare, or employee monitoring sits in a different regime with different obligations, and the correct advice is legal review before stage three, not a faster sequence. The Northwind illustration flags its CV-screening row and then leaves it alone, which is the right instinct.
Employment effects are outside its scope and genuinely unsettled. McKinsey's respondents split 32 percent expecting workforce decreases in the coming year, 43 percent no change, and 13 percent increases. Anyone who states the answer confidently is telling you their position, not a finding.
Where LeapForce Fits
LeapForce builds the control layer the last four artifacts live in: one governed endpoint to every model, identity in front of every AI surface with owner, scope and expiry for agents as well as people, a vetted connector registry with action-level scoping and human-in-the-loop gates, and tracing plus action audit that records what was refused as well as what ran. If you have completed the sequence manually, you have discovered the shape of that control layer by hand. The product is the version you do not have to maintain in spreadsheets. Our rollout guidance is the same as the sequence's: observe first, enforce second, optimise third. Per our published build status, gateway endpoints, tracing and SSO are live today, while capabilities such as vaulted keys, inline data-loss prevention and dollar budgets are in development and shadow-AI discovery and compliance evidence packs are on the roadmap — which is why stage one in this article is fifteen conversations rather than a scan.
What LeapForce does not do is choose your use case, run your discovery interviews, or replace a person's judgement at the commit gate. Stage one is human work, and it should be.
Frequently asked questions
No, and that has not been the binding constraint for a while. Assistants and no-code workflow builders let a non-technical operator assemble a trigger-plus-instruction workflow in an afternoon. What you do need is authority: someone able to approve a credential and an integration with your mailbox or CRM. In practice teams get stuck on permissions far more often than on syntax, which is why stage two of the sequence above is a decision rather than a build.
The one people in your company are already doing with AI on their own accounts, most often. Ask fifteen team leads what they used an AI tool for last month and sort the answers by frequency; take the most frequent item with the narrowest data footprint. This beats a prioritisation workshop for two reasons. A task somebody already chose to automate has passed a real test. And MIT's research found personal-tool use at over 90 percent of surveyed companies against 40 percent with an official subscription, so the working examples are already there.
Time savings on a single workflow are usually visible within two to four weeks of go-live, because you can count the runs. Getting to go-live is the slower part: eight to twelve weeks for a first cycle in a company with a normal IT function, dominated by waiting for credentials and system permissions rather than by building. Subsequent workflows take days, since the route and the cost mechanism already exist. Enterprise-level financial impact is a different timescale entirely; McKinsey found only 39 percent of organisations report any EBIT impact from AI at all.
For one workflow, model calls are usually the smallest line. In the worked example above, 550 runs a month cost about six dollars in inference and about $207 in human review time at the approval gate — human attention was 97 percent of the total. Budget accordingly: the honest cost of an AI workflow is inference plus review time plus the engineering hours to connect it, and the first of those three is rarely the one that matters. Read your own cost per run off your route's logs rather than a price list, because real prompts carry context and retries.
Usually yes for reading, and more carefully for writing. Mainstream systems of record have APIs and most workflow platforms have connectors, so extraction and drafting are straightforward. The write path is where design matters: restrict the workflow to a specific record status and a named list of fields, and keep the irreversible action, such as sending or pricing, behind a gate. A connector that can update anything in your CRM is a permission problem rather than an integration success.
It depends almost entirely on how narrowly you scoped it, not on which model you chose. OWASP's Top 10 for LLM Applications 2025 identifies Excessive Agency as a top risk and traces it to excessive functionality, excessive permissions and excessive autonomy — all three are choices made by whoever configures the workflow. Enumerate what the agent may read and write, run it in the user's own security context rather than under a shared privileged account, and enforce authorisation in the downstream system rather than in the prompt.
That depends on whether you gated the irreversible step. If the model's output only ever becomes a draft or a queued proposal, a wrong answer costs someone thirty seconds of correction and shows up as a row in your approval log. If the output commits directly, a wrong answer is an outbound email or an overwritten record. Design for the first case: assume a persistent error rate, put the gate where the action cannot be undone, and record refusals as well as approvals so you can see what the control caught.
Blocking without providing an alternative reliably moves the activity to phones and home laptops where you cannot see it, so sequence it the other way: build the approved route first, migrate the known uses onto it, then restrict. There is a real risk to address — one July 2026 survey found 38 percent of US workers had entered work information into a personal AI account their employer does not control — but the fix is a better path, not a firewall rule. Note also that 40 percent of respondents in that survey said their employer had no written AI policy at all, so a plainly worded policy is worth having alongside the route.
A named individual in the team whose work it does, with a named deputy, and not a committee or a department. The owner answers for output quality, holds the scope, and is the person a review date lands on. Assigning ownership to "Operations" or to IT means nobody answers on the day the output is wrong. Ownership also has to survive turnover: if the workflow was built by one person on their own account, promoting it into a company-owned asset with an explicit owner is what stops their resignation becoming an outage.
Use the first workflow's cost line as the unit and the approval log as the evidence. You can state runs per month, acceptance rate, correction rate, cost per run and review minutes per run — five real numbers from one workflow beat any projection. Then argue the marginal case: workflows two through six reuse the route and the cost mechanism, so their incremental cost is scoping and gate design, typically days rather than months. Be careful about promising efficiency alone; McKinsey found the organisations seeing the most value set growth or innovation as objectives too, not just cost reduction.
You need one thing before stage one, and it is narrower than a policy: a written statement that nobody will be penalised for disclosing how they have been using AI this week. Without that, your discovery returns a flattering and useless list. A fuller policy is worth writing after stage two, when you can point at an approved route and say what to use instead of the personal account. Writing the policy first tends to produce a document that forbids the only AI use your company currently has.
Often more clearly worth it, because the sequence runs faster: the owner is usually also the approver, so the permission waits that dominate mid-sized companies disappear and the first cycle can finish in three to five weeks. The stage that suffers is the cost line, because volumes are too small for a month of runs to say much. Run the same order anyway, keep the artifacts to a single page each, and accept that stage five will be an estimate for a while.
Ready to Govern Your AI?
Talk to LeapForce — one controlled layer for every AI tool, connector, model, and agent.
Comments