Invoice Automation: Price the Exception, Not the Scan

Invoice automation costs about the same per invoice as it did two years ago: an average of $9.84, according to Ardent Partners' State of ePayables 2025. The rea

Invoice automation costs about the same per invoice as it did two years ago: an average of $9.84, according to Ardent Partners' State of ePayables 2025. The reason is not that the software fails to read invoices. It reads them fine, for about three cents a page. The cost sits in the 18.4% of invoices that stop.

Our position is that almost every invoice automation business case is built on the wrong number. Buyers compare a blended cost per invoice before and after, and vendors quote a touchless rate. Neither tells you what you are buying, because the blended average is dominated by a minority of documents that never flow: the exception queue. Price that queue and the decision changes — including which product you should buy, and whether you should buy one at all.

The failure this is really about is not a typing error. On Hacker News, the practitioner johann8384 described a city manager who wired $150,000 abroad on the strength of a phishing email, and noted that the organisation's own rules "required an invoice and approvals" — the control existed and the payment went round it. Every serious project in this category is a decision about who, or what, is allowed to go round it next.

The short answer: Buy invoice automation on your exception rate and your approval thresholds, not on extraction accuracy or a touchless-rate promise — under a transparent cost model built on Ardent Partners' 2025 benchmarks, roughly 50% to 68% of what an average accounts payable function spends per invoice is consumed by the fifth of invoices that stop.

Last updated: July 30, 2026.

We have not run an invoice automation deployment ourselves and we are not going to pretend otherwise: every number below is either fetched from a named source or derived in the open from those sources, with the assumptions written down so you can change them.

Three lanes of invoice flow with modelled cost per invoice and each lane's share of total accounts payable spend

The three-lane split. Lane volumes are Ardent Partners' 2025 all-company benchmarks; lane prices are our derivation from them.

What invoice automation is, and the four things it actually does

Invoice automation is the set of software controls that take a supplier invoice from arrival to posted-and-approved without a person retyping it: capture, extraction, validation and matching, approval routing, and the ledger write. It is usually sold as accounts payable automation or automated invoice processing, and the useful way to think about it is not as one product but as four distinct jobs that fail in four different ways.

Job one is capture. Getting the document off an email, a portal, a scanner or an e-invoicing network into a system where it has an identifier and a timestamp. This is close to solved and nobody's business case turns on it.

Job two is extraction. Turning the picture of an invoice into fields: supplier, invoice number, date, currency, line items, tax, total. This is the part that vendors demo, the part that machine learning improved dramatically. As the pricing below shows, it is also the cheapest part of the whole chain.

Job three is decision. Two decisions, actually, and conflating them is the most common design error in accounts payable automation. Validation asks whether the extracted data is internally coherent: does it parse, do the lines sum, is the currency one you trade in, have you seen this invoice number from this supplier before. Matching asks whether it agrees with the purchase order and the goods receipt. Policy asks something else entirely — whether a coherent, matched invoice is one you are willing to pay without a person looking at it. A perfectly extracted invoice from a supplier nobody onboarded fails policy while passing validation.

Job four is commitment. Posting to the ledger and releasing the payment. This is the only irreversible step, it is where every control that matters lives, and it is the one an invoice automation demo spends the least time on.

What invoice automation is not: it is not a replacement for a purchase order process. If 35% of your invoices have no PO to match against, and Ardent's 2025 benchmark puts PO-linked invoices at 65.4% of the average company's volume, then automation cannot match what does not exist. It also is not the same thing as e-invoicing, which is about the format and the transport, not the decision. And it is not a fraud control by default; whether it becomes one depends entirely on what you do with the exceptions.

The 2026 benchmark numbers, including the one that moved backwards

The best public benchmark set for invoice processing cost comes from Ardent Partners, which has run an annual accounts payable study since 2006. It is also the only one that lets you check whether invoice automation is moving the metrics its vendors say it moves. Here is the 2025 edition's all-company table, taken directly from The State of ePayables 2025: AP's Unfinished Journey (June 2025), alongside the prior year's figures published in Ardent's AP Metrics that Matter in 2025 eBook, which compiled the 2024 survey of 212 accounts payable professionals.

Metric2024 survey2025 surveyDirection
Cost to process a single invoice (all-inclusive)$9.40$9.84Up 4.7%
Time to process a single invoice9.2 days8.2 daysDown 1.0 day
Invoice exception rate14.0%18.4%Up 4.4 points
Invoices processed straight-through32.6%35.4%Up 2.8 points
Suppliers submitting invoices electronicallynot published in eBook57.4%
Invoices linked to a purchase ordernot published in eBook65.4%
Staff time spent on supplier inquiries21.8%21.9%Flat

Two things in that table deserve more attention than they normally get.

The exception rate went the wrong way. The 2024-survey eBook celebrated exceptions dropping "dramatically to 14%." The 2025 survey puts them at 18.4%. In the same period, Ardent estimated that around three-quarters of accounts payable departments were already using some form of AI, and its 2025 report found 44% actively using AI with a further 34% planning to within twelve months. Whatever else the first wave of AI in accounts payable did, it did not shrink the exception queue.

Ardent's own trend labels do not reconcile with its own prior-year numbers. The 2025 benchmark table marks both cost per invoice and exception rate as "Declining," against published prior-year figures of $9.40 and 14.0%. We are not going to guess which figure is the corrected one. We report both, with their sources and dates, and flag the discrepancy — because a business case that quietly picks the flattering number from a pair like this is the reason so many of them do not survive their first review.

The maturity split matters more than the average anyway. Ardent defines Best-in-Class as the 20% of enterprises with the lowest per-invoice cost and shortest cycle times.

MetricBest-in-ClassAll Others
Cost to process a single invoice$2.65$12.42
Time to process a single invoice2.9 days13.5 days
Invoice exception rate11.1%20.9%
Invoices processed straight-through51.0%29.0%
Suppliers submitting invoices electronically67.2%47.3%
Invoices linked to a purchase order84.0%47.3%
Staff time on supplier inquiries12.8%24.0%

A $9.77 per-invoice gap between the top fifth and everyone else is the whole prize. What the table does not tell you is where inside the process that gap sits — and that is the question the rest of this piece answers.

One source we could not verify. APQC's cross-industry accounts payable cost benchmarks are the other set widely quoted in this category. Its measure pages returned 403 to every fetch method we tried. The APQC figures we could reach through secondary reporting, a $5.83 median against a $2.07 top quartile and $10-plus bottom quartile across roughly 1,485 organisations, all trace back to a 2018 article. We have left them out of the arithmetic rather than dress an eight-year-old benchmark as current. If you see those three numbers in a 2026 vendor deck, that is where they came from.

The three-lane split: why cost per invoice hides the decision

Every invoice you receive ends up in exactly one of three lanes, and a blended cost per invoice is the weighted average of three prices that differ by more than an order of magnitude. This is the split an invoice automation business case almost never makes.

  • The clear lane. Straight-through, touchless, no human hands. Ardent's 2025 average: 35.4% of volume.
  • The touched lane. Nothing is wrong; a person is nonetheless in the path — coding, a routine approval click, a portal download, a missing PO reference typed in. This lane is what is left over: 46.2% on the 2025 averages.
  • The exception lane. The invoice stops. Something does not match, does not sum, does not exist in the vendor master, or does not have an owner. Ardent's 2025 average: 18.4%.

We are assuming those three sets are disjoint — that an invoice recorded as an exception was not also recorded as straight-through. Ardent does not state this explicitly, and if some overlap exists the touched lane shrinks and the conclusions get stronger, not weaker.

The reason this split matters is that a vendor's headline promise almost always addresses the smallest cost lane. "We will take you to 60% touchless" is a promise about moving volume from the touched lane to the clear lane. That is real money. It is not where most of the money is.

Think of it as a queue with three tills. The blended figure tells you what the average customer paid across all three. It tells you nothing about which till has the line out of the door.

What an exception really costs

Nobody publishes a per-exception cost, which is awkward given that exception handling is the line item invoice automation is supposed to shrink. So derive it. Take Ardent's published averages as the constraint, add two assumptions you can argue with, and solve.

The constraint, for the all-company 2025 figures: 0.354 × (clear cost) + 0.462 × (touched cost) + 0.184 × (exception cost) = $9.84.

The two assumptions:

  1. A clear invoice costs about $1.00. Software licence, infrastructure and a share of overhead, amortised per document. It is not zero. It is also not much.
  2. An exception costs a multiple of a touched invoice. An exception involves at least one email to a supplier or a budget holder, a wait, and a second handling. We model three multiples — 3×, 4× and 6× — and show all three, because we cannot measure yours.
ScenarioClear invoiceTouched invoiceExceptionException lane's share of total AP cost
Exception = 3× touched$1.00$9.36$28.0752.5%
Exception = 4× touched$1.00$7.92$31.6759.2%
Exception = 6× touched$1.00$6.06$36.3568.0%

Read the last column again. Across every plausible assumption, between half and two-thirds of what an average accounts payable function spends per invoice is consumed by 18.4% of its documents. The clear lane, at 35.4% of volume, accounts for 3.6% of the spend.

The obvious objection is that the $1.00 clear-invoice figure is doing all the work. It is not. Hold the 4× multiple and vary the clear cost instead: at $0.50 the exception lane takes 60.3% of the blended cost, at $1.00 it takes 59.2%, and even at $3.00 — an implausibly expensive touchless invoice — it still takes 54.8%. The exception multiple moves the answer; the clear-lane assumption barely does.

Run the same model against the maturity split, holding the 4× multiple constant:

PopulationClearTouchedExceptionBlendedException lane share
Best-in-Class (51.0% / 37.9% / 11.1%)$1.00$2.60$10.40$2.6543.6%
All Others (29.0% / 50.1% / 20.9%)$1.00$9.07$36.29$12.4261.1%

The Best-in-Class advantage is not mainly that more invoices flow untouched. It is that when one stops, it costs $10.40 to resolve instead of $36.29 — and it stops half as often. Those two effects together explain the bulk of a 79% cost gap that gets casually attributed to "automation."

In money a CFO recognises, at 10,000 invoices a month:

UnitBest-in-ClassAll companiesAll Others
Per invoice$2.65$9.84$12.42
Per 1,000 invoices$2,650$9,840$12,420
Per month (10,000 invoices)$26,500$98,400$124,200
Per year (120,000 invoices)$318,000$1,180,800$1,490,400
Exception volume per year13,32022,08025,080
Annual exception-lane cost (4× model)$138,528$699,274$910,153

Scale from the per-1,000 row if your volume is different; the ratios do not change with size, only the totals do. The line that should reframe a procurement conversation is the last one. An average 120,000-invoice-a-year operation is spending roughly $700,000 annually on twenty-two thousand documents that stopped. A tool that improves extraction accuracy by three points does not touch that number. A tool that cuts the exception rate from 18.4% to 11.1% removes about 8,760 exceptions a year, and at $31.67 each that is roughly $277,000 — before you count what happens to the resolution cost itself.

Extraction is the cheapest part of invoice automation

Here is the number that ends the "AI is expensive" objection and, less comfortably for vendors, the "AI is the value" claim at the same time. Google Cloud's published Document AI pricing, fetched 30 July 2026:

ProcessorPublished pricePer page
Enterprise Document OCR (1,000–5,000,000 pages/month)$1.50 per 1,000$0.0015
Enterprise Document OCR (above 5,000,000/month)$0.60 per 1,000$0.0006
Custom extractor (up to 1,000,000/month)$30.00 per 1,000$0.030
Custom extractor (above 1,000,000/month)$20.00 per 1,000$0.020
Layout Parser$10.00 per 1,000$0.010

Three cents a page for a trained custom extractor. On a $9.84 blended cost that is 0.3%. Across 120,000 invoices a year it is $3,600 — about 0.3% of the $1.18m that operation spends processing them, and roughly one two-hundredth of what the same operation spends on its exception queue.

That has two consequences a buyer should carry into every demo.

Token and inference spend is not the thing to negotiate. It is real, it belongs in the model, and on any realistic volume it is rounding. If a vendor's pricing conversation is mostly about per-document AI charges, they are drawing your attention to the cheap end of the process.

Extraction accuracy is also not the thing to negotiate, for the same reason inverted: it is nearly free to run, so the marginal value of the last accuracy point is small compared with the cost of what happens when it is wrong. The next section is about exactly that.

Two honest caveats. These are list prices for one cloud provider's document processing service, not the price of a packaged product, which bundles workflow, ERP connectors, supplier onboarding and support — the per-document component is a small part of what a vendor charges. And a general-purpose multimodal model called through an API prices differently again, on tokens rather than pages, with image tokenisation making a scanned page more expensive than these figures. The order of magnitude is what carries: cents, not dollars.

The accuracy claim that does not survive contact with a ledger

"Over 99% extraction accuracy" appears in almost every invoice automation page in this category, almost always with no source and no definition of what is being measured. Two 2026 pieces of published research let us put a floor under the reality.

The strongest published result we could find on real invoices is from "Information Extraction from Electricity Invoices with General-Purpose Large Language Models", submitted to arXiv on 1 April 2026. On a subset of the IDSEM dataset of Spanish electricity invoices, the best configuration — few-shot prompting with cross-validation — reached an F1 score of 97.61% with Gemini and 96.11% with Mistral-small. The authors also report that the gap between zero-shot and the best few-shot strategy "exceeds 19 percentage points," which is to say that most of the accuracy in a well-tuned extraction pipeline comes from work someone did on the prompting and the examples, not from the model.

Note the conditions: one country, one industry, one broadly consistent invoice layout, with tuning. That is the easy end of the problem. The harder end is measured by UNIKIE-BENCH, a February 2026 benchmark (revised April 2026) that evaluated 15 state-of-the-art large multimodal models on key information extraction from visual documents. Its finding, in the authors' words, is "substantial performance degradation under diverse schema definitions, long-tail key fields, and complex layouts, along with pronounced performance disparities across different document types and scenarios."

Now do the arithmetic that vendor pages skip. Field-level accuracy is not document-level accuracy. An invoice is not correct because most of its fields are correct; it is correct when all of them are.

If a document has 12 fields that matter and each is independently right 97.61% of the time, the probability that the whole document is right is 0.9761 to the twelfth power — about 74.8%. Roughly one invoice in four would carry at least one wrong field.

Fields per invoiceField accuracy 97.61%Field accuracy 99.0%Field accuracy 99.5%
686.5%94.1%97.0%
1274.8%88.6%94.2%
2061.7%81.8%90.5%
40 (line-item heavy)38.0%66.9%81.8%

The independence assumption is doing real work there and it is wrong in both directions. Errors correlate: a bad scan or an unusual layout breaks many fields at once, which makes the true document-level rate better than the table implies because failures cluster. On the other hand, line-item-heavy invoices have far more than 12 fields, which makes it worse. Treat the table as a shape, not a measurement.

But the shape is the point, and it reconciles with something we already have. Ardent's observed exception rate of 18.4% is the same order of magnitude as the 25% document-level error implied by a 97.6% field accuracy across 12 fields. Extraction quality alone plausibly accounts for a large share of the exception queue, which is why the vendor claim of 99%-plus deserves one specific question: 99% of what — fields, or documents? If the answer is fields, ask for the document-level straight-through rate on your own invoice mix during a pilot, on a sample you chose.

The Lane Audit: a one-month diagnostic before you buy anything

You do not need a consultant to find your own numbers, and the exercise takes one accounts payable calendar month. Run it before you shortlist a single invoice automation vendor. We call it the Lane Audit because its whole job is to sort your volume into the three lanes above and put a price on each.

Prerequisites — you need read access to one month of invoice records with timestamps, a way to see who touched each one, and thirty minutes a week from whoever runs the queue. You do not need a tool, a pilot, or a vendor.

You almost certainly do not need every invoice. A random sample of 200 gives you the lane split to within roughly plus or minus five percentage points at 95% confidence — the standard error on a proportion near 20% at n=200 is the square root of (0.184 x 0.816 / 200), or 2.7 points, so the interval is about 5.4 points wide either side. That is precise enough to tell a 12% exception rate from a 25% one, which is the decision you are making. Sample randomly across the whole month rather than taking the first 200, because the first week of a month and the last week are different processes.

The seven columns. One row per invoice, or per invoice for a representative sample if your volume is large:

ColumnWhat goes in itWhy
Invoice IDYour own referenceJoins back to the ledger
LaneClear / Touched / ExceptionThe whole point
Stop reasonFree text, one lineBecomes the fix list
Human touchesCount of distinct peopleDrives cost
Elapsed daysReceipt to postedDrives discount capture
Value bandUnder $1k / $1k–10k / $10k–100k / over $100kDrives the threshold design
PO linkedYes / NoPredicts the exception rate

What the month usually reveals. Three things, in our reading of the benchmark data and the way the stop reasons cluster in the field. First, the stop reasons concentrate: a handful of causes will account for most of the exception lane, and the top one is very often "no purchase order" rather than anything a model could have read better. Second, the human-touch count is higher than anyone guessed, because approval chasing does not feel like processing. Third, your exception rate and your PO-linkage rate move together — Ardent's Best-in-Class have 84.0% of invoices PO-linked against 47.3% for everyone else, and half the exception rate.

Pricing the lanes. Take your total accounts payable running cost for the month, including salaries, benefits, software, overhead and any outsourced processing. Divide by invoice count for your blended figure, then apply the same two assumptions we used above and solve for your own three lane prices. If your blended number is far from $9.84, that is information, not an error.

The output is one sentence you should be able to say out loud before any vendor call: "We process N invoices a month at $X blended; Y% stop; the top three stop reasons are A, B and C; and the exception lane costs us $Z a year." A vendor who cannot tell you which of A, B and C their product removes is selling you the clear lane. That sentence, built from your own month, is what goes into the paper you take upstairs — not our model, which exists only to tell you where to look.

If a month is too long, run two weeks and double it, accepting that you will miss month-end effects — which is precisely when the exception queue is worst, so treat the result as a floor.

Approval authority: which invoices software may clear on its own

Everything else in invoice automation hangs off this one decision, and it is a business decision wearing a technical costume. It belongs in the approval workflow, not in the model. The question is not "can the system approve invoices" but "which invoices, at what value, from which suppliers, with what evidence, and who signed off on that rule."

We use a four-question test, in order. An invoice may clear without a human only if the answer to all four is yes.

  1. Is the counterparty already trusted? Approved supplier, in the vendor master, with banking details that have not changed inside your cooling-off window.
  2. Is there something to check it against? A purchase order and, for goods, a receipt. No PO, no auto-clear — this single rule is why PO coverage predicts exception rate.
  3. Is it inside tolerance? Both a percentage and an absolute cap, because 2% of $2m is not a rounding error.
  4. Is it below the value band where a person must own the decision? Set by your own distribution, not by a vendor's demo.

That last one needs its own paragraph, because it is where most implementations quietly go wrong. Pick the band by opening your own invoice-value histogram and asking what share of value, not volume, falls below each candidate threshold. In most accounts payable populations the distribution is heavily skewed: a large majority of documents carry a small minority of the money. That skew is the friend of a well-designed threshold, because a band that clears most of your volume can still leave nearly all of your money in front of a human.

Choose the low band (auto-clear only under a few hundred dollars) if you are in your first two quarters, if your vendor master has not been cleaned, or if your auditors have flagged accounts payable in the last cycle. You will get a smaller efficiency win and almost no new risk.

Choose the middle band (auto-clear PO-matched, in-tolerance invoices up to a few thousand) if PO coverage is above roughly 80% and your exception stop reasons are concentrated and understood. This is where most of the realisable saving is.

Choose the high band, or no band at all, so that every invoice sees a human before payment. Do this if you are in a regulated flow, if a single wrong payment would be materially damaging, or if the supplier population changes often. Some organisations should be here, permanently, and no amount of touchless-rate benchmarking should move them.

Two rules that hold at every band. Confidence is not authority: a model's confidence score is a routing signal, never a permission. Let it decide which queue an invoice enters; never let it decide whether money leaves. And a timeout is not an approval: if nobody clicks, the run waits. Our earlier analysis of human-in-the-loop automation sets out the tests that separate a real approval gate from a rubber stamp, including why a gate a reviewer can never realistically say no to is not a control at all.

Segregation of duties when the approver is software

The oldest control in accounts payable is that the person who sets up a supplier is not the person who approves the invoice is not the person who releases the payment. The GAO's Green Book (GAO-25-107721, Standards for Internal Control in the Federal Government) states the principle plainly. It binds US federal agencies rather than private companies, and it is adapted from the COSO internal control framework that private auditors do work from; we quote it because it is the clearest statement of the principle that anyone can read for free, not because it applies to you. Paragraph 10.22: "Segregation of duties helps prevent fraud, waste, and abuse in the internal control system," and management "considers the need to separate control activities related to authority, custody, and accounting of operations."

Invoice automation breaks this quietly, in a way that no dashboard shows you. When one integration user account performs extraction, matching, posting and payment release, every one of those separated duties has been recombined inside a single identity — and the ledger now says "AP_INTEGRATION" for all of it. You have not removed the control; you have removed the evidence that it was ever there.

The Green Book anticipates the case where separation is not practical. Paragraph 10.23: if segregation of duties "is not practical within a business process because of limited personnel or other factors, management designs alternative control activities to mitigate the risk." That is the standard to design to, and it gives you the shape of the answer.

Three moves, in ascending order of effort.

Split the identities. The extraction and matching job runs under a credential that can read purchase orders, read the vendor master and write nothing. The posting job runs under a second credential that can write journal entries and cannot modify the vendor master. Payment release is a third. If your ERP cannot express that, that constraint belongs in your evaluation criteria, above the extraction accuracy line.

Give every automated identity an owner, a scope and an expiry. A service account with no named human owner is the thing that survives three reorganisations and gets discovered during an audit. We have written separately on treating non-human identities as first-class: owner, scope, expiry, and an offboarding path. It is the failure mode that generalises well beyond accounts payable.

Keep vendor-master maintenance outside the automated path entirely. Changing a supplier's bank details is the single highest-value action in the whole process and it should never be reachable by the same identity that approves invoices, however convenient that would be.

AP Now, a specialist accounts payable education channel, covers the manual version of these controls in depth, including three-way match and separation of duties, in this April 2026 session. It is useful if you need to explain to a non-finance stakeholder why the control existed long before software was involved.

Play video

The exceptions that are not errors: duplicates, bank details, fraud

Some exceptions are data problems. Some are somebody trying to get paid twice, and some are somebody trying to get paid who is not your supplier at all. An invoice automation design that routes both kinds into the same "needs review" bucket is the design that eventually pays the wrong account.

The scale is not speculative. The Association for Financial Professionals' 2026 Payments Fraud and Control Survey, conducted in January 2026 among 465 treasury practitioners, found that 76% of US organisations experienced attempted or actual payments fraud in 2025, that 74% were affected by business email compromise, and, the line that matters most here, that just 17% of organisations use AI to combat payments fraud. AI has been adopted far faster on the processing side of accounts payable than on the control side.

The FBI's 2025 Internet Crime Report puts a figure on the consequence: business email compromise accounted for $3,046,598,558 in reported losses in 2025, up from $2,770,151,146 in 2024, against 1,008,597 total complaints and $20.877 billion in overall reported losses, a 26% year-over-year increase. BEC is, structurally, an invoice-and-approval attack.

And the insider version is slower and more expensive. The ACFE's Occupational Fraud 2026: A Report to the Nations, covering 2,402 cases across 143 countries, reports a median loss of $104,000 per case and a median scheme duration of 12 months. Schemes caught inside six months cost a median $40,000; those running over five years cost more than $1.1 million. Detection speed is most of the loss. Tips remain the leading detection method at 43% of cases, which is a polite way of saying that in most organisations the controls did not catch it.

So separate your exception queue by kind, not just by severity:

Exception kindExampleCorrect handling
Data defectLines do not sum; tax field unreadableAuto-retry, then a person; safe to fix in-line
Reference gapNo PO, no receipt, unknown cost centreRoute to the requester, never resolve inside AP
Policy breachUnapproved supplier, over tolerance, over bandNamed approver; log the override reason
Integrity signalDuplicate candidate, changed bank details, new payee first invoice, unusual urgencyEscalate out of the AP queue; second-channel verification; never auto-resolve

That fourth row is the one automation most often gets wrong, because a duplicate-invoice check that silently suppresses the second copy looks exactly like good automation and destroys the signal. A duplicate that arrives from a different email domain three weeks later is not a data-quality event.

A rule worth writing down: an automated system may detect an integrity signal and it may route one, but it should never clear one. The exit from that queue is a human decision recorded against a named person, with the verification channel noted.

The per-invoice record an auditor will accept

Most invoice automation projects log the outcome. Very few log the input, and the gap only shows up six months later when someone asks why a particular payment went out.

Here is the per-invoice record we would design to. Each line answers a question somebody eventually asks.

FieldAnswers
Run ID and receipt timestampWhen did this arrive, and which processing attempt is this?
Source and senderWhere did the document come from — which mailbox, portal, or network?
Document hashIs this byte-for-byte the file we processed?
Extracted values, as extractedWhat did the machine actually read?
Corrected values and who corrected themWhat changed between reading and posting, and on whose authority?
Model and version identifierWhich system produced the extraction, so a later defect can be scoped
Validation and match resultsWhich checks ran, and what did each return?
Policy version appliedWhich rules were in force that day, not which are in force now
Decision and deciderAuto-cleared under rule R, or approved by named person P
Refusals and overridesWhat the system declined to do, and who overrode it
Ledger referenceWhich journal entry, keyed idempotently to the run ID

Two fields on that list are the ones almost nobody has and both are cheap to add at build time and impossible to reconstruct afterwards.

The policy version. An auditor asking about an August payment wants the rules as they stood in August. A system that can only show the current threshold cannot answer.

The refusals. A log that records only what happened cannot demonstrate that a control was working; it only demonstrates that nothing blocked. Recording what the system declined to do, and what a human overrode, is what turns a log into evidence. We have set out the broader argument for audit trails that prove agent actions, including why refusals belong in the record.

Two practical notes. Retention: the record has to outlive the audit that will ask for it, which in most jurisdictions means the same retention period as the underlying accounting records rather than your application log's default of ninety days — decide this at build time, because the extracted values are unrecoverable once purged. Regulatory scope: ordinary accounts payable processing is not a high-risk use under the EU AI Act, so the Act's logging obligations are unlikely to bite here directly; our guide to EU AI Act compliance for deployers sets out which uses do fall in scope and on what timetable, if you are unsure where your process sits.

And keep the idempotency requirement in view: the ledger write should be keyed on the run ID so that a retry after a network failure cannot post the same invoice twice. That is not an AI feature. It is the reason the AI feature is safe to use, and our earlier analysis of where to cut a business workflow between rules and models works through the same boundary across nine back-office processes.

What the 2026 e-invoicing mandates do to the business case

If you operate in Europe, part of this decision has been made for you, and the timing is now close enough to change what you should buy.

France's reform is the nearest hard date. Per the French tax administration's own guidance on when the electronic invoicing obligation applies, from 1 September 2026 every company in scope must be able to receive invoices in electronic form, while large and mid-sized enterprises must also issue them electronically and transmit transaction and payment data to the administration. Small, medium and micro-enterprises pick up the issuing obligation from 1 September 2027. Reception is universal on the first date; issuing phases in.

Three consequences for an invoice automation business case written in 2026.

The capture problem partly solves itself, on somebody else's timetable. Structured e-invoices arrive as data, not as pictures of data. For the share of your volume that arrives that way, extraction accuracy stops being a question at all — which further undermines the case for choosing a product on extraction quality.

Your exception mix changes rather than shrinking. Structured invoices remove reading errors. They do not create purchase orders, do not clean a vendor master, and do not decide whether a matched invoice is one you want to pay. Expect the exception lane to become proportionally more policy and reference gaps and less data defect — which is exactly the mix that needs approval design rather than better models.

Buy for the transport you will actually be on. A product priced on per-document capture is a product whose value declines as e-invoicing coverage rises. Ardent's 2025 figure already puts electronic submission at 57.4% of invoices on average and 67.2% among Best-in-Class, and 83% of accounts payable leaders expected the count to rise that year.

If you are outside Europe, this still matters when you have European suppliers or subsidiaries, and it is a reasonable forward indicator for other jurisdictions.

The costed middle path, and when invoice automation does not pay

The choice is not automate-everything or do-nothing, and the middle option is usually the one that survives a business case review. Here is what each looks like, costed against the same 120,000-invoice-a-year operation and the 4× exception model.

OptionWhat you actually doModelled annual costWhat it does not fix
Do nothingStatus quo at the all-others benchmark$1,490,400Everything
Fix the inputs onlyPO discipline, vendor-master cleanup, supplier e-invoicing enablement. No new platformFalls between $318,000 and $1,490,400 depending on how far exception rate and PO linkage moveApproval routing; the touched lane stays
Clear-lane automationAutomate capture, extraction, matching and posting for PO-matched in-tolerance invoices only; every exception and every above-band invoice stays with peopleClear-lane invoices approach the $1.00 marginal cost; the exception lane is unchangedThe exception lane, which is 50–68% of your spend
Full platformAbove, plus exception workflow, supplier portal, approval routingBest-in-Class benchmark $318,000 is the ceiling of the ambition, not the quoted priceNothing structurally — but it is the largest change-management load

The three levers that actually move the exception rate, in order of cost to pull: purchase-order coverage, because a missing PO is the most common stop reason and no model can invent one; vendor-master hygiene, because unapproved and duplicate supplier records generate policy exceptions that look like data errors; and supplier e-invoicing enablement, because structured input removes reading errors at source. None of the three is a software purchase.

The honest ranking for most mid-market organisations is: fix the inputs first, because it is the cheapest lever and it improves the return on everything you buy afterwards; then automate the clear lane, because it is low-risk and self-evidently correct; then take on exception workflow, which is where the money is and also where the design decisions above become unavoidable.

When invoice automation does not pay. Three cases, stated plainly.

Volume below the fixed cost. Every platform carries an implementation and subscription floor. It does not scale down. Below a few hundred invoices a month, the arithmetic often favours cleaning up the process and leaving it manual. Run the Lane Audit before you assume otherwise; your blended cost may be well under the benchmark simply because your process is small and short.

A supplier population that changes constantly. Project-based construction, media production and similar patterns generate a long tail of one-time suppliers. Matching has nothing to match against and the exception lane is the process, not the deviation. Automation here optimises the smallest lane.

No purchase order discipline. If most invoices arrive with no PO, no automation product can validate them — it can only route them faster to the same person who was going to have to decide anyway. Fix the front of the process first.

Where the governance layer sits

Once the invoice decision is made by software rather than a person, the questions that decide whether it is safe stop being accounts payable questions and become deployment questions: what identity did the agent act under, what was it allowed to touch, which policy version was in force, what did it refuse, what did it cost, and who owns it when its builder leaves. That layer is what LeapForce builds: access, policy, cost and audit, applied uniformly across every AI tool, connector, model and agent. LeapForce does not sell an invoice automation product, an accounts payable module or an ERP connector for three-way matching, and you should not evaluate us against the vendors in that category.

What we would carry over from our own rollout model is the ordering. Our gateway rollout guide is "Observe first. Enforce second. Optimize third," and it maps onto this problem almost exactly: run one month of the Lane Audit in observe mode before any rule is written, then set the approval bands and identity scopes, and only then tune for touchless rate. Teams that invert that order buy a straight-through percentage and discover their controls afterwards. Per LeapForce's published build-status convention, capabilities on our platform are labelled LIVE, IN DEV or ROADMAP; check the platform pages for the current status of any specific capability before you plan around it.

Where this analysis is uncertain

Several things here are modelled, not measured, and you should know which.

The lane prices are a derivation, not a benchmark. Ardent publishes lane volumes and a blended cost. The split into clear, touched and exception prices is ours, resting on two stated assumptions — a $1.00 clear invoice and an exception costing 3–6× a touched one. We showed three scenarios because we cannot measure yours. If your clear invoice costs $3 rather than $1, every figure moves; the qualitative finding, that the exception lane dominates, survives across the range we tested but is not immune to a very different assumption set.

The disjoint-lane assumption is unstated in the source. We assume an invoice recorded as an exception was not also counted as straight-through. Ardent does not say. If they overlap, our touched lane is too big.

The two Ardent editions disagree with themselves. Exception rate reads 14.0% for the 2024 survey and 18.4% for the 2025 survey, and the 2025 table labels that trend "Declining." We have used both figures with their sources rather than reconcile them, and a reader who needs one number should go to the primary reports rather than to us.

The document-level accuracy table assumes independent field errors. They are not independent. The table is a shape and we have said so where it appears.

Ardent's respondent base is self-selected. These are accounts payable professionals who answered an industry survey; that population plausibly skews toward larger and more engaged functions. Sample sizes are 212 for the 2024 edition; the 2025 edition does not state one we could verify.

We could not verify APQC's current figures. Its benchmark pages returned 403 to every method we tried, and the accessible reporting of its per-invoice benchmarks dates to 2018. Those numbers are excluded from our arithmetic for that reason.

Reddit was unreachable for this research, so the practitioner voice here comes from Hacker News, which skews technical for a topic whose primary audience is finance. Treat the single field voice we cite as illustrative rather than representative.

No first-hand deployment. We have not implemented invoice automation in a live accounts payable function and have not measured a real exception rate ourselves. Everything above is derived from named public sources and stated assumptions.

 FAQ

Frequently asked questions

Invoice automation is software that takes a supplier invoice from arrival to posted and approved without anyone retyping it. It covers five steps — capture, data extraction, validation and matching against a purchase order and goods receipt, approval routing, and the ledger write. It is also called accounts payable automation or automated invoice processing. The important distinction is between validation, which asks whether the extracted data is internally coherent, and policy, which asks whether a coherent invoice is one you are willing to pay without a human. Most implementations conflate the two and inherit the exception problem that follows.

The all-inclusive cost of processing an invoice covers staff, software, overhead and outsourcing. It averaged $9.84 in Ardent Partners' State of ePayables 2025, against $2.65 for the top-performing fifth of organisations and $12.42 for everyone else. Those are total processing costs, not software prices. The machine-reading component is trivially small: Google Cloud's published Document AI pricing puts a trained custom extractor at $30 per 1,000 pages, or three cents per invoice page, fetched 30 July 2026. If you want your own number, divide your total monthly accounts payable running cost by your invoice count before you take any vendor's word for the baseline.

Ardent's 2025 all-company average is 35.4% of invoices processed straight-through, rising to 51.0% for Best-in-Class organisations and 29.0% for everyone else. A vendor quoting 80% or 90% touchless is quoting a ceiling achievable on a clean, PO-linked, single-supplier subset, not an average across a real invoice mix. The best predictor of where you will land is your purchase-order linkage rate: Best-in-Class organisations have 84.0% of invoices linked to a PO against 47.3% for the rest, and half the exception rate.

On real invoices in favourable conditions, the best published 2026 result we found is an F1 score of 97.61% using Gemini with few-shot prompting and cross-validation, on Spanish electricity invoices from the IDSEM dataset. On harder material, the UNIKIE-BENCH benchmark of 15 multimodal models reports substantial degradation with diverse schemas, long-tail fields and complex layouts. Critically, those are field-level scores. If a document has 12 fields each right 97.61% of the time and errors were independent, only about 75% of documents would be entirely correct. Always ask a vendor whether their accuracy claim is per field or per document.

An exception is any invoice that stops before it can be posted. Ardent's 2025 average exception rate is 18.4%, up from 14.0% in the prior year's survey — though the 2025 report labels that trend "Declining," a discrepancy worth knowing about. Exceptions come in four kinds that need different handling: data defects, reference gaps such as a missing purchase order, policy breaches such as an unapproved supplier or an over-tolerance amount, and integrity signals such as duplicates or changed bank details. Only the first is safe to fix inside the accounts payable queue.

Only those passing four tests together: the supplier is already approved with unchanged banking details, there is a purchase order and receipt to match against, the variance is inside both a percentage and an absolute tolerance, and the value falls below a band you set from your own invoice-value distribution. Set the band by asking what share of value, not volume, sits beneath each candidate threshold — in most accounts payable populations a band that clears most documents still leaves nearly all the money in front of a person. Two rules never bend: a model's confidence score is a routing signal rather than a permission, and a reviewer timeout is never an approval.

We cannot give you a credible universal figure and we would distrust anyone who does, because the variable that dominates is not the software but the state of your purchase-order discipline and vendor master. What you can do in one month, before signing anything, is the Lane Audit described above: sort a month of invoices into clear, touched and exception lanes, price each, and identify your top three stop reasons. That produces a payback estimate specific to you, and it tells you whether the product you are considering removes your stop reasons or only the ones in its own demo.

Every serious product integrates with the major accounting and ERP systems, so the answer is almost always yes and it is the wrong question. The better questions are what the integration can write and under which credential. Ask whether the connector can be scoped so that extraction and matching read purchase orders and the vendor master without write access, whether posting runs under a separate identity, and whether vendor-master maintenance can be excluded from the automation path entirely. If those are not separable in your ERP, that constraint should rank above extraction accuracy in your evaluation.

Not a shared integration account performing every step. Split the credentials so extraction and matching can read but not write, posting can write journal entries but not modify supplier records, and payment release is separate again — which is the GAO Green Book's segregation-of-duties principle applied to a non-human actor. Where separation genuinely is not practical, the Green Book's own guidance at paragraph 10.23 is to design alternative control activities, which in practice means tighter logging and a compensating human review. Every automated identity should also carry a named owner, an explicit scope and an expiry date.

At minimum: run identifier and receipt timestamp, source and sender, a hash of the document, the values as extracted, any corrected values with who corrected them, the model and version that produced the extraction, validation and match results, the policy version in force at the time, the decision and who or what made it, any refusals and overrides, and the ledger reference keyed idempotently to the run ID. The two most commonly missing are the policy version and the refusals. Without the policy version you cannot answer what the rules were in the month being audited; without refusals you can only show that nothing blocked, not that a control was working.

If you operate in France, yes and soon. Per the French tax administration, from 1 September 2026 all in-scope companies must be able to receive electronic invoices, while large and mid-sized enterprises must also issue them and transmit transaction and payment data; small and micro-enterprises take on the issuing obligation from 1 September 2027. The practical effect is that structured invoices arrive as data, so the capture-and-extraction part of the problem shrinks on a timetable you do not control. That argues against paying a premium for extraction quality and in favour of products whose value sits in exception workflow, matching and approval design.

Often not as a platform purchase, and the reason is fixed cost rather than capability. Implementation and subscription floors do not scale down, so below a few hundred invoices a month the arithmetic frequently favours fixing the process — purchase-order discipline, a clean vendor master, getting suppliers onto electronic submission — and leaving the handling manual. Two exceptions are worth checking: if your exception rate is unusually high because of poor supplier data, the cheap fix is the data rather than the software; and if you are in scope for an e-invoicing mandate, receiving structured invoices becomes an obligation regardless of your volume.

Ready to Govern Your AI?

Talk to LeapForce — one controlled layer for every AI tool, connector, model, and agent.

Thirty minutes · No pitch deck

Ready to turn AI experiments into measurable ROI?

Bring one outcome you'd like AI to move. We'll help you scope a pilot you can actually measure — and tell you honestly if it's not worth doing yet.

Comments