AI Process Optimization: Fix the Constraint, Not the Task

AI process optimization is the practice of using models to find where a business process loses time or accuracy, and then changing that specific point — not the

AI process optimization is the practice of using models to find where a business process loses time or accuracy, and then changing that specific point — not the whole process, and not the loudest part of it. It works when the change lands on the step that actually governs the outcome. It fails, most of the time, because nobody measured which step that was.

Our position is narrower than the usual pitch, and it is arithmetic rather than opinion: the maximum improvement any intervention can produce is capped by the share of total lead time the targeted step consumes. Automate a step worth 3% of lead time perfectly and you have bought a 3% improvement. That buys you a model, an integration, a review queue and a new failure mode. On Hacker News in December 2025, a developer posting as asdev asked whether anyone was running document processing with language models in production for ERP data entry, and put the trap plainly: "the human would just have to read the document themselves to ensure accuracy" (Ask HN, 18 December 2025). They were describing a real process-optimization failure in advance — work moved from one queue to another, with the total unchanged.

The short answer: Measure lead time per step before you buy anything; the only step worth optimizing with AI is the one whose wait time and rework rate dominate the process, and improvement is mathematically capped at that step's share of total lead time.

Last updated: July 30, 2026.

We have not run a controlled AI process optimization program inside a customer's finance department and measured the before-and-after ourselves, so nothing below is presented as our measurement. The numbers in the worked example are an explicitly modelled process, labelled as such; every external figure carries its source inline.

Stacked bar showing an invoice process where waiting time dwarfs touch time, with two candidate AI interventions marked

The anatomy of a five-step invoice process, modelled: 147 hours of waiting around 37 minutes of work. Where you point the model decides everything.

What AI Process Optimization Actually Is

AI process optimization is the use of machine learning and language models to observe how a process runs, identify where it loses time or accuracy, and change that point — either by handling a step directly, by routing work differently, or by surfacing a problem earlier than a human would notice it. The word doing the work in that sentence is observe. Everything downstream depends on whether the observation was real.

It is worth saying plainly what it is not, because four things get sold under this name and only one of them is the thing.

Often sold as AI process optimizationWhat it actually isWhen it is genuinely useful
A chatbot on top of an existing systemInterface changeWhen the interface was the barrier, which is rare
Task automation for one repetitive taskSubstitution of manual work at one pointWhen that step is the constraint, or when the step is the whole process
A dashboard that describes the processMeasurement, not optimizationAlways — but it is the input, not the outcome
Continuous adjustment of routing, thresholds or assignment based on observed outcomesOptimization properWhen you can tell what changed and reverse it

The fourth row is the interesting one, and it is also where the honest caveats live. A system that continuously adjusts how work flows is a system that is quietly changing your operating procedure. That is a governance question long before it is a technology question, and we come back to it below.

The distinction most articles draw, that rules do what they are told while AI adapts, is true but oversold. Rule-based automation is not a failure of imagination. It is a deliberate trade of adaptability for determinism, and in a lot of processes determinism is the thing you are actually paying for. We wrote about that trade in our earlier analysis of why AI business automation starts with rules rather than models, and the argument holds here: the question is not rules or models, it is which parts of a process reward each.

The Measurement Gap: Why AI Lands on the Wrong Step

Most organizations cannot say, to the nearest hour, how long their core processes take end to end, or which step is the real process bottleneck. They can tell you how many invoices they processed. They cannot tell you how long the median invoice waited between arriving and being matched. That gap is the reason AI process optimization so often produces an impressive demo and a flat operating result.

The academic discipline built to close that gap is process mining, which reconstructs the real process from the event logs your systems already write. The Process and Data Science group at RWTH Aachen University frames one of its three core capabilities as conformance checking — answering, in their words, whether people "really follow the expected process." The premise embedded in that question is the important part: the documented process and the executed process are different objects, and the executed one is the only one AI can improve.

The macro data says the same thing from a different angle. In the US Census Bureau's Business Trends and Outlook Survey, national AI use among businesses stood at 19.8% as of the collection period ending 3 May 2026, rising to 37% among firms with 250 or more employees. But the Census research team's own working paper on how that adoption is distributed found something sharper. In The Microstructure of AI Diffusion (Bonney, Breaux, Dinlersoz, Foster, Haltiwanger and Pande, April 2026), the authors report that "worker task use sometimes occurs without formal firm-level adoption, and firm-level adoption sometimes occurs without worker task use", with 57% of adopting firms deploying AI across three or fewer business functions and 65% restricting use to three or fewer tasks.

Read that as an operations finding rather than a statistics finding. AI is arriving at the task layer, chosen by whoever felt the pain, and the firm layer often does not know. Nobody in that arrangement is looking at the process end to end, which means nobody is looking at where the process is actually slow. The optimization target gets selected by proximity to frustration, not by contribution to cycle time.

Stanford's 2026 AI Index describes the same shape at the enterprise level: 88% of surveyed organizations had adopted AI, generative AI is used in at least one business function at 70% of organizations, and yet AI agent deployment sits "in the single digits across nearly all business functions." Broad adoption, shallow integration. Tools layered onto processes rather than processes rebuilt around tools.

The three questions that expose the gap in one meeting

You do not need a software purchase to find out whether your organization has this problem. Ask three questions about the process you were about to optimize, and time how long it takes to get an answer.

QuestionIf the answer takes more than a dayWhat it tells you
What is the median end-to-end lead time for this process, last quarter?You have no baseline, so you cannot detect an improvementAny post-launch claim of improvement will be unfalsifiable
Which single step holds an item longest, counting waiting as well as working?You are about to pick a target by anecdoteThe intervention will probably land off-constraint
What percentage of items come back for rework, and from which step?Quality cost is invisible to youAI that increases throughput will increase rework proportionally

Every AI process optimization program we have seen described publicly that reports a disappointing outcome is consistent with failing at least one of these three. That is a pattern claim, not a measurement, and we mark it as such.

The Ceiling Rule: The Arithmetic That Decides Where AI Pays

Here is the rule the rest of this article rests on. The maximum improvement available from perfecting a step is that step's share of total lead time. Not its share of effort, not its share of headcount, not its share of complaints. Its share of the clock.

This is not a new insight; it is Amdahl's law wearing a business suit, and Eliyahu Goldratt made the operational version famous in the theory of constraints. What is new is how easily AI lets you violate it. A model can automate a step in an afternoon. That speed removes the friction that used to force a team to ask whether the step was worth automating.

Work the arithmetic on a modelled process. The figures below describe a five-step supplier-invoice flow at a mid-sized company handling roughly 1,000 invoices a month. These are illustrative figures constructed to show the method, not measurements from a customer. The shape — a small amount of touch time surrounded by a large amount of waiting — is the standard shape of an office process, and the point survives substantial changes to the numbers.

StepTouch time per invoiceWait before the stepRework rateShare of lead time
1. Receive and capture6 min4 h12%2.8%
2. Match to PO and receipt3 min2 h5%1.4%
3. Exception resolution25 min71 h18%48.4%
4. Approval2 min46 h3%31.2%
5. Payment run1 min24 h1%16.3%
Total37 min147 h100%

Total lead time is roughly 147.6 hours, about 6.2 calendar days. Total touch time is 37 minutes. The ratio of the two, value-added time over lead time, is the lean measure called process cycle efficiency, and here it is 0.42%. Ninety-nine and a half percent of an invoice's life is spent waiting for someone.

Now apply the ceiling rule to the two interventions a vendor demo would most likely propose.

InterventionWhat it targetsCeiling on lead-time improvementRealistic result
Document model extracts invoice fields automaticallyStep 1 (2.8% of lead time)2.8%~1.4%, since capture still waits on the inbox cycle
Model triages exceptions on arrival, classifies the cause, drafts the supplier query and routes itStep 3 (48.4% of lead time)48.4%~35% if the 71-hour wait falls to 20 hours

The first intervention is the one that demos beautifully. It is visible, it is impressive, the before-and-after screenshot is compelling, and it caps out at a 2.8% improvement to the thing the business actually cares about. The second is duller to watch and worth roughly twenty-five times more.

This is also exactly the failure the Hacker News poster anticipated. Automating capture without touching exception resolution moves the work into a verification queue. If the model is right 95% of the time and a human has to check all of it to find the 5%, the review step inherits most of the original touch time and the process gains almost nothing. A commenter in the same thread, posting as muzani, replied from real experience running auto insurance claim extraction and made the counterpoint that matters: "Semi-auto beats manual readily," because their design flagged mismatches automatically against a second source rather than asking a person to re-read everything. The difference between those two outcomes is not model quality. It is whether the design removed the verification burden or relocated it.

The Constraint-First Loop

We use a four-move loop for this, and we name it so it travels: Baseline, Constraint, One Change, Ledger. It is deliberately smaller than the seven-step implementation plans that circulate on vendor blogs, because the failure mode is not insufficient planning. It is starting at step four.

Four-move cycle diagram showing Baseline, Constraint, One Change and Ledger, with the ledger feeding back into the baseline

The Constraint-First Loop. The fourth move is the one that makes the loop a loop rather than a project.

Move 1 — Baseline

Pick one process. Instrument four columns per step: touch time, wait time before the step, rework rate, and volume. That is the whole baseline. Keep it crude. A crude baseline collected this month beats a precise one scheduled for next quarter.

Where the numbers come from, in descending order of quality: system timestamps you already have (ticket created, ticket assigned, ticket resolved); an export from the system of record with state-transition history; a two-week sample where a few people log start and stop times; and last, informed estimates from the people doing the work, explicitly marked as estimates. Estimates are legitimate — they are just labelled differently, and you should expect them to be wrong about waiting time in particular, because waiting is invisible to the person waited on.

Two things to get right at this stage. First, measure the median and the 90th percentile, not the mean. Office processes have long tails, and the tail is usually where the pain and the escalations live. Second, count waiting as time. Every handoff between people or systems creates a queue, and queues are where cycle time is spent. The single most common baselining error is recording only the minutes someone was actively working, which produces a picture in which the process takes 37 minutes and everyone wonders why customers complain about the six days.

Move 2 — Constraint

The constraint is the step with the largest share of lead time, adjusted for rework. It is usually not the step people complain about. People complain about steps that are annoying to perform, and the process bottleneck is typically a step where nothing is happening at all.

Three diagnostics find it quickly.

DiagnosticHow to run itWhat it reveals
Lead-time shareRank steps by wait plus touch time, as in the table aboveThe ceiling on any intervention at each step
Queue depth over timeCount items sitting in each state, daily, for two weeksWhere work accumulates rather than flows; a queue that grows is a constraint, a queue that empties nightly is not
Rework originFor each item that came back, record which step produced the defectWhere speeding up would multiply downstream cost

Queue depth deserves a note. Little's Law, the queueing result that average work-in-progress equals average arrival rate multiplied by average time in system, is why counting the pile is such a cheap proxy for measuring the clock. If sixty invoices are permanently sitting in exception resolution and you clear twenty a day, items are spending three days there, and you learned that from a count rather than from instrumentation.

Rework origin is the diagnostic people skip and later regret. If step 1 produces defects that surface at step 3, accelerating step 1 accelerates defect production. The DORA research team put the general version of this well in their 2025 report on AI-assisted software development: AI adoption showed a positive relationship with throughput and a negative relationship with software delivery stability, a pairing the report summarizes as acceleration exposing weaknesses downstream when control systems are missing.

Move 3 — One Change

Change one thing, at the constraint, with a written hypothesis, a stop rule and a review date. The hypothesis has a number in it. "AI-assisted exception triage will reduce median wait at step 3 from 71 hours to under 30 hours within six weeks" is a hypothesis. "AI will make invoice processing more efficient" is a press release.

The stop rule is the part teams leave out and it is the part that makes the loop safe to run. Write down, in advance, the condition under which you turn the change off. Rework rate at the following step rises above its baseline. Cost per item exceeds a stated ceiling. Median lead time does not move by the review date. Anything measurable and pre-committed.

One change at a time is not caution for its own sake. If you deploy three interventions simultaneously and lead time drops 20%, you have learned that something worked, which is not a transferable finding. The next process gets no benefit from it.

Move 4 — Ledger

Every change to how the process runs gets a row. This is the move that distinguishes an optimization capability from an optimization project, and it is the one that vendor implementation guides almost universally omit.

Ledger fieldWhy it is there
What changed, in one sentenceSix months later nobody remembers
Which process and stepSo the effect can be attributed to a baseline
Who approved itAccountability for an operating-procedure change
What identity and scope it runs underWhich systems the change can touch, and with what permissions
Hypothesis and stop ruleMakes the result falsifiable in advance
Measured effect at the review dateCloses the loop, or documents that it did not work

A ledger row costs about four minutes to write. When the process misbehaves in November and someone asks what changed, the ledger answers in seconds. Without it, the honest answer is that between March and November roughly eleven things changed, three of them automatically, and nobody can order them.

The loop then restarts: the ledger's measured effect becomes the new baseline, and you find the next constraint, which will be a different step, because relieving a constraint always moves it.

What AI Genuinely Changes About a Constraint

Once you know where the constraint is, the question becomes whether AI is the right instrument for that particular constraint. Sometimes it is not, and buying it anyway is how organizations end up with a governed, observable, expensive system doing something a calendar rule would have done.

Four capabilities are genuinely new, in the sense that no previous automation technology offered them at reasonable cost.

CapabilityThe constraint it relievesThe failure mode to watch
Reading unstructured inputSteps where a human must interpret an email, PDF, call transcript or free-text field before work can startConfident extraction of a wrong value, which is worse than an obvious failure
Classification and triage at arrivalQueues where items wait because nobody has decided who should handle themMiscategorised items that land in a slow queue and are never escalated
Continuous watchingProblems detected on a weekly report that could have been detected on arrivalAlert volume that exceeds anyone's attention, restoring the original delay
Drafting a first versionSteps where the cost is composition, not decisionReviewers rubber-stamping drafts because reviewing feels like agreeing

Three of the four attack waiting time rather than working time, which is precisely why they matter in processes shaped like the table above. Triage-at-arrival returns the most of the four in a typical office process, and it is the least demoed, because a well-triaged queue looks like nothing happening.

Note what is missing from that list. AI does not create approval capacity when the constraint is a named human who must sign. It does not remove a step that exists because a regulator requires it. It does not resolve a queue owned by a department that has not agreed to change. Those are political and structural constraints. Applying a model to them produces a faster arrival at the same wall. Before committing, check whether the constraint you found is a work constraint or a decision constraint. We set out the general test in our analysis of when approval is the control rather than the bottleneck.

That leaves an obvious question for the reader whose constraint turns out to be one of those: what do you do instead? The honest answer is that the highest-return fix for a decision constraint is usually free, and the reason teams miss it is that nobody was looking at lead time.

If the constraint isThe usual AI answerThe fix that actually works
Managers batching approvals once or twice a weekDraft the approval summary fasterChange the batching cadence, or auto-approve under a threshold with sampling
A queue owned by another departmentRoute items to it fasterNegotiate a service-level target; the constraint is an agreement, not a technology
A step that exists because of a policy nobody has revisitedAutomate the stepDelete the step
One named approver who is the single point of sign-offSummarise their inboxAdd a second delegated approver with defined limits
Items waiting on an external party's responseDraft the chase messageChange the contract terms or the chase schedule

Run this test before writing a budget line. If the constraint appears in the right-hand column, an AI project will produce a faster arrival at the same wall, and the ledger will faithfully record that nothing improved.

Where the returns actually concentrate

The most useful finding in the productivity literature is not the headline average, it is the distribution around it. Brynjolfsson, Li and Raymond's study of a generative AI assistant deployed to customer support agents found productivity, measured as issues resolved per hour, rose 14% on average — but with novice and lower-skilled agents improving 34% while experienced, highly skilled workers saw minimal gains.

That is the ceiling rule again in a different coordinate system. The intervention paid where the gap between current performance and achievable performance was largest. Applied to processes rather than people, the same logic says: look for the step with the widest spread between its best-case and typical execution, not the step with the highest volume.

Stanford's 2026 AI Index reports gains in the same range for the same setting, 14% to 15% in customer support, 26% in software development and 50% in marketing output. Those are task-level measurements from studies of specific functions, and they should not be read as what a process will yield end to end. A 26% improvement in the coding step of a release process that spends most of its lead time in review and approval moves the release date very little — which is, once more, the ceiling rule.

When the Rigid Rule Still Wins

The standard comparison table in articles on this topic sets "traditional automation" against "AI-powered optimization" and lets the second win every row. That table is a straw man, and enterprise readers know it, which is part of why the category has a credibility problem. Rule-based automation wins on several axes that matter enormously in production.

PropertyDeterministic rulesModel-driven optimization
Same input, same outputGuaranteedNot guaranteed without pinning and caching
Testable before deploymentFully, with a finite test setStatistically, on a sample
Cost of diagnosing a failureRead the ruleReconstruct the context, the prompt, the version and the inputs
Marginal cost per executionEffectively zeroMetered per call, and it scales with volume
Handles unanticipated inputFails visiblyProduces something plausible, which may be worse
Auditor's question "why did it do that?"Answerable by inspectionAnswerable only from logs you decided to keep in advance

Practitioners reach the same conclusion from the other direction. In the same Hacker News thread, a commenter posting as ensemblehq described doing document extraction for capital markets and medical assessments by "starting with something rules-based and working through a layered multi-model setup", after a pure-model attempt on medical documents produced poor output. That is not a rejection of models. It is a statement about where in the stack each belongs.

Row five is the one that decides more architectures than any other. A rule that encounters an unexpected input throws an error, and an error is a gift: it is loud, it is diagnosable, and it stops the process at a known point. A model that encounters an unexpected input returns a confident answer. In a payments process, the second behaviour is strictly worse than the first.

The synthesis that survives contact with production is not "AI replaces rules". It is AI proposes, deterministic pipeline executes. The model reads the messy input, classifies it, and produces a structured proposal. A rule engine validates that proposal against constraints that cannot be violated — this vendor is on the approved list, this amount is under this threshold, this account code exists. The rule engine executes. The model never touches the commit.

That split gives you the adaptability where you needed it, at the interpretation step, and determinism where you needed it, at the point of consequence. It is also the design that makes audit tractable, because the deterministic half is where the irreversible action lives. We have argued the general form of this elsewhere: in AI-powered workflows, the step to govern is the commit, not the builder or the model.

There is a third case worth naming, which the "when the incumbent wins" framing usually misses: sometimes the right answer is to delete the step. A process with a 71-hour exception queue may have a 71-hour exception queue because of a policy written eleven years ago for a supplier relationship that ended in 2021. Optimizing it with a model is expensive nostalgia. The Baseline move surfaces these, and the correct outcome of a process-optimization exercise is occasionally that you cancel the project and remove a rule.

What the Evidence Actually Shows About Returns

Anyone proposing a budget for AI process optimization is walking into a room where somebody has read a headline about failure rates. It is worth having the actual numbers, including the unflattering ones.

SourceFindingWhat it does and does not support
MIT NANDA, The GenAI Divide: State of AI in Business 2025 (Challapally, Pease, Raskar and Chari, July 2025)"95% of organizations are getting zero return" despite $30–40 billion of enterprise investment; method was 52 structured interviews plus analysis of 300+ public AI initiativesSupports scepticism of pilots; a 52-organization interview base is not a representative sample, and "zero return" partly reflects the absence of measurement
DORA / Google Cloud, 2025 State of AI-assisted Software Development90% of ~5,000 respondents use AI at work; over 80% report increased productivity; AI adoption is positively related to throughput and negatively related to delivery stabilityStrong evidence that speed and stability move in opposite directions without control systems; software-specific
Brynjolfsson, Li and Raymond, Generative AI at Work (NBER w31161)14% average productivity gain in customer support; 34% for novices; minimal for expertsThe best available causal estimate for one function; not generalisable to all processes
Stanford HAI, 2026 AI Index, Economy chapter88% of organizations adopted AI; generative AI in at least one function at 70%; AI agent deployment in the single digits across nearly all functionsAdoption is near-universal, autonomous deployment is not; the gap is where governance work sits
US Census Bureau BTOS, collection ending 3 May 202619.8% national AI use; 37% among firms with 250+ employees; under 20% among firms with fewer than 20The most representative adoption number available for US businesses; a firm-level self-report, not a usage measurement
Census CES-WP-26-25, The Microstructure of AI Diffusion57% of adopting firms use AI in three or fewer functions; task use occurs without firm adoption and vice versaDirect evidence that adoption is task-shaped rather than process-shaped

The most quoted of these is the 95% figure, and it is worth handling honestly rather than either dismissing or weaponising it. The MIT NANDA report's finding is about measurable P&L impact from pilots. A large part of that 95% is not projects that failed to change anything; it is projects that changed something nobody had baselined, run by teams who could not therefore prove it. That is not an argument that the pilots worked. It is an argument that the first deliverable of an AI process optimization program is a measurement system, and that programs which skip it forfeit the ability to defend themselves later. Our own reading of why programs stall at the pilot boundary is set out in AI pilot to production.

The DORA framing is the one we would put in front of an executive sponsor. Their conclusion is that AI acts as an amplifier — strong teams get stronger, struggling teams find existing problems intensified. Applied to processes: a well-instrumented process gets faster, and a badly instrumented one gets faster at producing whatever it was already producing, including the defects.

Two sources a reader would reasonably expect here are missing, and their absence is deliberate rather than an oversight. Reddit threads on this topic are inaccessible to automated retrieval, so the practitioner voices below come from Hacker News, which skews technical relative to the finance and operations audience this topic actually serves. McKinsey's State of AI survey, widely quoted in this category, could not be retrieved to verify the figures attributed to it, so it is not cited.

Here is a recorded lecture on conformance checking from the Process Mining Summer School, run by the academic community that built the discipline. It is long and technical. It is also the clearest available explanation of how you establish what your process actually does before you change it.

Play video

A Worked Example, End to End

This section assembles the whole method on the modelled invoice process introduced above. Again: the numbers are constructed to demonstrate the arithmetic, not measured at a customer.

Move 1, Baseline. Two weeks of state-transition timestamps pulled from the accounting system, plus a rework tally kept by the AP team on a shared sheet. Result: the table earlier in this article. Median lead time 6.2 days, 90th percentile 14 days, process cycle efficiency 0.42%, rework concentrated at step 3 (18%) and step 1 (12%).

Move 2, Constraint. Exception resolution, at 48.4% of lead time. Queue depth confirms it: the exception queue holds roughly 60 invoices at any time and clears about 20 a day, implying three days resident, which lines up with the 71-hour wait measured from timestamps. Approval at 31.2% is the runner-up, but the diagnosis there is different — approval waits because managers batch their sign-offs, which is a behaviour problem with a scheduling fix, not a model problem.

Move 3, One Change. Hypothesis: an AI triage step that reads the exception, classifies it into one of six known causes, drafts the supplier or requester query, and routes it to the named owner of that cause will reduce median wait at step 3 from 71 hours to under 30 hours within six weeks, without increasing the downstream rework rate at approval.

Stop rules, written before launch:

Stop conditionThresholdAction
Misclassification rateAbove 10% on a weekly 30-item auditPause, retune categories
Approval-stage reworkAbove 5% (baseline 3%)Pause, investigate defect origin
Cost per exception handledAbove $0.40Route to a cheaper model or cap volume
Median wait at step 3Not below 45 h by week 4Stop and reconsider the constraint

The design decision that matters. The model does not resolve exceptions. It classifies, drafts and routes. A human still decides, and a rule still validates the eventual correction against the approval matrix. This is the AI-proposes-pipeline-executes split, and it is what keeps the intervention from converting a resolution queue into a verification queue — the exact failure the Hacker News poster predicted.

Expected result if the hypothesis holds. Wait at step 3 falls from 71 hours to 20. Total lead time falls from 147.6 hours to 96.6 hours, or 6.2 days to 4.0 — a 35% reduction, against a ceiling of 48.4%. Touch time barely moves, which is why an effort-based business case would have missed this entirely and a cycle-time-based one catches it.

Move 4, Ledger. One row, written the day the change goes live.

FieldEntry
ChangeAI triage classifies AP exceptions into six causes, drafts the query, routes to cause owner
Process and stepSupplier invoice to payment, step 3, exception resolution
Approved byAP manager, with finance controller sign-off on the approval-matrix interaction
Identity and scopeService identity with read access to the invoice store and write access only to the exception queue; no payment permissions
Hypothesis and stop ruleMedian wait 71 h to under 30 h by week 6; stop rules as tabled above
Measured effect at reviewTo be completed at week 6 against the same query used for the baseline

Two details in that ledger are easy to skim past and are the ones that matter in an incident. The identity row says the triage agent cannot touch payments — so when someone asks in month four whether the optimizer could have caused a duplicate payment, the answer is a scope definition rather than an investigation. And the measured-effect row specifies the same query used for the baseline, which closes the most common way an improvement claim gets quietly manufactured: measuring the after with a different definition than the before.

The Half Nobody Writes: Who Approved the Change

Continuous AI process optimization means a system is changing how work happens while you are not watching. That is the entire value proposition, and it is also a description of an unlogged change to a business process. Every other article on this keyword treats the governance question as a security footnote about data. The harder question is operational: what changed, who allowed it, what was it permitted to touch, and can you put it back.

This is the layer LeapForce builds. Not process mining, and not the optimizer itself — we do not sell a process discovery tool and would not pretend to. What we build is the control layer underneath: a governed gateway every model call passes through, non-human identities so an optimizer runs as itself with an owner, a scope and an expiry rather than borrowing a person's credentials, connector scoping so the triage agent in the example above can read invoices and write to one queue and nothing else, budgets denominated in dollars rather than tokens so an always-on optimizer cannot quietly outspend the savings it found, and an audit trail that records refused actions alongside executed ones. Our rollout model for that layer is the same order this article argues for the process itself — observe first, enforce second, optimize third, as set out on our AI gateway page: point traffic at the gateway in observation mode, learn what is actually happening, and only then write rules. Per-capability build status is published honestly on the site, and some of the observability surface is still in development rather than shipping today.

Regulation is converging on the same requirement from the outside. The EU AI Act's Article 12 obliges high-risk systems to "technically allow for the automatic recording of events (logs) over the lifetime of the system", so that system functioning is traceable. If your process optimizer sits inside a process that touches employment decisions or creditworthiness, that logging is a legal obligation rather than an engineering preference. Even where it does not, the NIST AI Risk Management Framework, released in January 2023, organizes its guidance around Govern, Map, Measure and Manage — a sequence in which measurement precedes management, which is the same claim this article makes about processes. We go further into what a defensible action record contains in our work on AI observability and audit trails.

The practical version, for a team that has no governance layer at all yet: the Ledger move is the minimum viable one. A shared document with the six fields above, filled in for every process change, gets you most of the accountability benefit at zero cost. Tooling helps when the number of changes exceeds what a person will diligently record, which happens sooner than teams expect.

Your First Thirty Days

A sequence for business process optimization that fits inside a normal quarter, assumes no new software purchase in the first month, and produces something falsifiable at the end.

DaysActivityOutput
1–3Pick one process with a clear start and end event and a business owner who will answer questionsA named process and a named owner
4–10Baseline: pull state-transition timestamps, count queue depth daily, tally reworkThe four-column table, with sources labelled by quality
11–13Identify the constraint by lead-time share, queue growth and rework originOne step, with its share of lead time stated
14–16Decide whether the constraint is a work constraint, a decision constraint or a policy artefactA go, a no-go, or a rule to delete
17–20Write the hypothesis, the stop rules and the review date; define the identity and scope the change will run underA one-page change proposal
21–25Build the smallest version that tests the hypothesis; keep the commit step deterministicA working intervention at one step
26–30Launch to a slice of volume, write the ledger row, schedule the reviewA live change and a row that will answer questions later

Two warnings about this schedule, both learned from how these programs fail rather than from running this exact plan ourselves. First, days 4 to 10 will take longer than you budgeted, every time, because the timestamps are worse than anyone believes. Budget the slip there rather than compressing the baseline, since a compressed baseline invalidates everything after it. Second, resist adding a second intervention during days 21 to 25 when the first one looks easy. The discipline of one change is what makes the result transferable to the next process.

Where This Is Still Uncertain

Several things in this article are weaker than the confident register might suggest, and it is better to say so than to have a reader discover it.

We did not run this. No LeapForce team member ran a controlled AI process optimization program in a customer's accounts payable function and measured the outcome. The invoice figures are a constructed model chosen to show arithmetic that holds across a wide range of inputs, and they are labelled as such at each appearance. If you want measured numbers for your own process, the Baseline move produces them in about a week, and they will differ from ours.

The ceiling rule is a ceiling, not a forecast. It tells you the maximum available from perfecting a step. It does not tell you what you will achieve, and a step's share of lead time can itself shift as volume, staffing and seasonality change. Re-baseline before making a second claim.

The evidence base is uneven across functions. The strongest causal estimate available, the 14% customer support finding, comes from one function at one company. DORA's data is software delivery. The Census figures are firm-level self-reports about adoption, not measurements of process outcomes. Nobody has yet published a large, credible study of AI's effect on end-to-end lead time in ordinary back-office processes, which is precisely the number this article would most like to cite.

The rework interaction is under-quantified everywhere. We know from DORA that throughput and stability can move in opposite directions. We do not have a general rule for how much rework a given accuracy rate produces in a given process, because it depends entirely on where the defect is caught. Your own rework-origin tally is more informative than any published figure.

Continuous optimization has a failure mode we are still learning to name. A system that adjusts routing weekly based on observed outcomes will eventually optimize for the metric you instrumented rather than the outcome you wanted, and it will do so gradually enough that no single week looks wrong. The stop rules and the ledger are our current answer. We do not think they are a complete answer.

When this is the wrong exercise entirely. If your process has low volume, no repeated shape, and lead time driven by a single external party's response speed, none of this applies. Optimizing a process that runs eleven times a year is a hobby. Spend the effort on the process that runs eleven thousand times.

 FAQ

Frequently asked questions

AI process optimization is using models to observe how a business process actually runs, find where it loses time or accuracy, and change that specific point. In practice it takes three forms: reading unstructured input so work can start sooner, classifying and routing items at arrival so they stop waiting for a human to decide who owns them, and watching continuously so problems surface on arrival rather than on a weekly report. The optimization part is the change; the AI part is what makes the observation cheap enough to run continuously.

Traditional business process optimization is periodic and human: map the process, find the bottleneck, redesign, re-measure next quarter. Robotic process automation replays a fixed sequence of clicks and does exactly that forever. AI process optimization differs on two axes — it can handle input that has no fixed structure, and observation can run continuously rather than in quarterly waves. What it does not change is the arithmetic: an adaptive system applied to a step worth 3% of lead time still yields at most 3%. The genuine advantage of rules is determinism, testability and near-zero marginal cost, which is why most production designs end up with the model proposing and a rule engine executing.

Pick the process with high volume, a clearly defined start and end event, timestamps you already have, and a business owner who will answer questions. Then, within that process, target the step with the largest share of total lead time after adjusting for rework — not the step people complain about most. Complaints track how unpleasant a step is to perform; lead time tracks what the business feels. In the modelled invoice process in this article, the annoying step was capture at 2.8% of lead time and the constraint was exception resolution at 48.4%.

Four columns per step: touch time, wait time before the step, rework rate, and volume. Sources in descending order of quality are system state-transition timestamps, an export with status history, a two-week manual sample, and informed estimates explicitly labelled as estimates. You do not need clean data and you do not need a warehouse. You need the median and the 90th percentile of end-to-end lead time, and you need waiting counted as time — the most common baselining error is recording only active work, which yields a process that appears to take 37 minutes while customers experience six days.

Not to start. Process mining tools automate the reconstruction of the real process from event logs and are genuinely valuable at scale or across many processes, but the first baseline of a single process is a spreadsheet exercise using timestamps your systems already write. Buy the tooling when the manual approach breaks — typically when you are running the loop on more than a handful of processes, or when the process spans several systems whose logs need stitching. The academic discipline behind the category, described by the Process and Data Science group at RWTH Aachen, frames the core question as conformance checking: whether people really follow the expected process.

Budget one week for the baseline that will actually take two, three days for constraint identification, and four to six weeks for one change to produce a signal you can defend. That puts a first defensible number somewhere between eight and ten weeks for a single process. Anything faster is usually a demo rather than a result, and anything measured before the process has cycled enough times to give a stable median is noise. Set the review date in advance and measure with the same query you used for the baseline.

Three cost lines, and most business cases forget the second and third. Model inference is metered per call and scales with volume, so a triage step on 1,000 items a month has a very different bill from one on 100,000. Human review time is often the larger line — if a change relocates verification rather than removing it, you have bought a cost without a saving. And there is the standing cost of the control layer: logging, identity, monitoring. Set a per-item cost ceiling as a stop rule before launch, and denominate budgets in currency rather than tokens so that finance can read them.

Show a baseline taken before the change, using a defined query, and the same query after. State the ceiling, meaning the constraint step's share of lead time, so the claim has an upper bound that is obviously not marketing. Report median and 90th-percentile cycle time, not the average. Include the costs above. And be prepared for the objection from the MIT NANDA finding that 95% of the organizations it studied were getting zero return on generative AI: the correct response is that a large share of that number reflects the absence of measurement rather than the absence of effect, which is exactly what a pre-committed baseline prevents.

Someone named, before it goes live, and recorded. Continuous optimization means a system is altering an operating procedure between reviews, so the approval question needs answering at the level of the change class rather than the individual change: this agent may adjust routing within these queues, may not alter approval thresholds, and may not act on items above this value without a human gate. Write the identity, the scope and the approver into a ledger row for every change. If a regulator, an auditor or a customer asks in six months what changed and who allowed it, the ledger is the answer; reconstructing it afterwards is expensive and usually incomplete.

It becomes one when the process is in scope for regulation and the changes are not recorded. The EU AI Act's Article 12 requires high-risk systems to automatically record events over the system's lifetime so that functioning is traceable, and processes touching employment or creditworthiness fall into that category. The NIST AI Risk Management Framework, published in January 2023, organises its guidance as Govern, Map, Measure and Manage. Neither prohibits adaptive optimization. Both assume you can say what the system did and why, which is a design requirement to build in at the start rather than retrofit after an audit request.

The baseline and constraint work require a spreadsheet and access to timestamps, not code. Building an intervention increasingly does not either — no-code and low-code builders now cover classification, routing and drafting steps competently. What still requires engineering judgment is the boundary: keeping the commit step deterministic, defining what identity the change runs under, and deciding what it may touch. Those are architecture decisions rather than programming tasks, but skipping them because the builder made the intervention easy is how a no-code optimization ends up with production write access nobody scoped.

A small team can run the loop, and in some ways more easily, because the process fits in one person's head and the timestamps live in fewer systems. The economics differ though: US Census data shows under 20% of firms with fewer than 20 employees using AI at all, against 37% of firms with 250 or more. For a small team the honest guidance is to run Baseline and Constraint manually, and to be ruthless about the ceiling rule — at low volume, the fixed cost of building and governing an intervention is a large fraction of the saving, so a constraint worth less than about a third of lead time is rarely worth automating. Deleting the step is disproportionately often the right answer at that scale.

Ready to Govern Your AI?

Talk to LeapForce — one controlled layer for every AI tool, connector, model, and agent.

Thirty minutes · No pitch deck

Ready to turn AI experiments into measurable ROI?

Bring one outcome you'd like AI to move. We'll help you scope a pilot you can actually measure — and tell you honestly if it's not worth doing yet.

Comments