AI Agent Business Applications: Who Signs the Work

AI agent business applications succeed in the functions whose sign-off already runs at machine speed. Before choosing a use case, name the person who countersig

AI agent business applications succeed in the functions whose sign-off already runs at machine speed. Before choosing a use case, name the person who countersigns that output today, measure how long an error takes to surface, and price what unwinding it costs after that delay.

Almost every guide to this topic sorts agents by what they can do: draft, triage, reconcile, route. We think that sorting hides the thing that actually decides whether the deployment survives contact with the business. An agent does not create a new accountability structure; it inherits the one the owning department already has. Finance closes the books once a month with a controller's signature on it. Support answers a customer within seconds and nobody signs that reply at all. Same model, same autonomy setting, same connector — completely different consequence when it is wrong. On Hacker News in April 2026, a user running fourteen agents in production put the gap plainly: the frameworks all solve how agents pass data to each other, but "Those are organizational problems, not technical ones" was their verdict on authority, disagreement and escalation (dsteel on Hacker News, 5 April 2026).

The short answer: Pick AI agent business applications by the countersignature the owning function already runs: who signs, how fast an error surfaces, what unwinding costs, and which rule attaches — because agents scale first where that signature was already automated, not where the task is easiest.

Last updated: July 31, 2026.

Six business functions plotted by how fast an agent error surfaces against what it costs to unwind

Where each function sits on detection speed versus reversal cost. The safe corner is bottom-left, and it is not where most pilots start.

What an AI Agent Business Application Actually Is

An AI agent business application is a named, scoped software worker that carries a multi-step task across more than one system, keeps context between steps, and acts without a human re-triggering each step. That is the capability definition, and it is the one every vendor page uses. It is also incomplete, because it describes the agent and says nothing about the department that will answer for its output.

The complete definition has a second half. An AI agent business application is a task the agent performs plus the accountability structure the owning function already operates. Both halves ship together whether you design the second one or not. The moment an agent posts a journal entry, that entry enters a close that a controller will sign. The moment an agent sends a customer an answer, that answer becomes a representation the company made. The agent did not create either obligation; it inherited one that predates it by decades.

What it is not: an AI agent business application is not a chatbot with integrations, and it is not a macro. A chatbot produces sentences a human then acts on, which keeps the human as the signer by construction. A macro produces deterministic output, so a single review of the rule covers every future run. An agent breaks both properties at once — it acts, and each run can differ from the last. That is precisely why the question of who signs stops being rhetorical.

Nor is it a department. "Sales agents" is a category, not an application. The unit that matters is one class of output: the outbound follow-up email, the invoice-to-PO match, the candidate rejection — because a class of output is the smallest thing a signature can attach to.

Where Agents Are Actually Running, By Function

Agent deployment is broad and shallow, and the shallowness is concentrated in a pattern worth reading carefully. In its State of AI survey published 5 November 2025, McKinsey found 23 percent of respondents scaling an agentic system somewhere in the enterprise and a further 39 percent experimenting — but "in any given business function, no more than 10 percent of respondents say their organizations are scaling AI agents." Most of those scaling are doing so in one or two functions only.

Which functions? McKinsey reports agent use is most commonly reported in IT and knowledge management, where service-desk management and deep research have developed fastest. That result is usually read as a technical-affinity story: engineers build agents, so engineering gets agents. We read it differently. IT is also the only function in a normal company that had already automated its own countersignature before agents arrived: code review, continuous integration, staged rollout, one-command rollback. The output of an IT agent lands in a place where a machine already checks machine-speed work.

Knowledge management fits the same pattern from the opposite direction. A research brief has no countersignature at all because it has no external effect until a human uses it. Nothing needs signing, so nothing blocks.

McKinsey State of AI, fielded 25 June – 29 July 2025Figure
Organizations using AI in at least one business function88%
Scaling an agentic AI system somewhere in the enterprise23%
Experimenting with AI agents39%
Scaling agents in any single given business functionno more than 10%
Reporting at least one negative consequence from AI use51%
Reporting consequences from AI inaccuracynearly one-third

The last two rows are the ones a department head should sit with. Half of organizations using AI have already had a negative consequence, and inaccuracy is the most commonly reported one. McKinsey also notes that the second-most-reported risk, explainability, is not among the most commonly mitigated. Explainability is exactly what a countersigner needs in order to sign. The gap between "we experienced this" and "we mitigate this" is, functionally, a gap in signatures.

We checked for a newer edition before using these figures; the November 2025 survey remains McKinsey's current State of AI release as of 31 July 2026.

The Countersignature Test: Four Fields Per Function

Run this before choosing a use case, not after. It takes one meeting per candidate output class and produces four fields. We call it the Countersignature Test, and each field is a question about the department, not about the agent.

Field 1. Signer. Who countersigns this class of output today, by name and role? Not "the team." A named human whose approval makes the output official. If the honest answer is "nobody signs each one, we sample," write that down; it is a valid answer and it changes everything downstream.

Field 2. Surface time. If this output is wrong, how does the error become visible, and how long does that take? Measure from the moment the output leaves the department to the moment somebody who can act on it knows. Use the department's own historical incidents to calibrate, not an estimate.

Field 3. Unwind cost. What does it cost to reverse the output after Field 2's delay has elapsed? Reversal cost is not constant over time, and this is the field most teams get wrong by pricing an instant undo that will never be available.

Field 4. Regime. Which external rule, contract, or professional duty attaches to this class of output? Employment law, consumer protection, financial reporting standards, bar admission, sector regulation. If a regulator, an auditor, a tribunal or a licensing body has an opinion about this output, name it here.

Score the four together, not separately. A short surface time with a high unwind cost is worse than a long surface time with a low one, which is the inversion most risk frameworks miss because they treat speed of detection as an unqualified good.

Four-field flow of the Countersignature Test with the two outcomes it produces

The Countersignature Test in sequence. Any field that comes back empty is a design task, not a blocker.

Why Fast Detection Is Not the Same as Low Risk

Detection speed and reversal cost are independent variables, and the interesting business applications sit where they disagree. A support agent's mistake is detected in seconds, because the customer reads it and replies, but the statement has already been made, and in the wrong circumstances it binds the company. A finance agent's mistake may sit undetected for weeks, but as long as it is caught inside the accounting period, reversing it is a journal entry.

That inversion is the single most useful output of the test. It explains why teams that start with customer-facing agents because "we'll see problems immediately" end up with the incident that reaches the executive team, while teams that start with reconciliation and matching, slow to detect and cheap to reverse, quietly get to production.

The general rule we now use: deploy where the unwind is cheap for the whole length of the detection window. Both halves matter. A cheap unwind that expires in an hour is not cheap if nobody looks for four days.

There is a third variable hiding inside Field 2 that deserves its own line: whether the error surfaces at all. Some outputs have no natural detector. A rejected job candidate does not tell you the rejection was wrong. A prospect who never replies to an inaccurate outbound email does not tell you why. Where the detector is missing, the department is not running an unmonitored process. It is running an unmonitorable one, and the agent will make that invisible failure faster and larger without changing its visibility.

Six Functions at a Glance

Six common homes for AI agent business applications, scored on the four fields. Detection latency and unwind cost are typical figures for the class of output named, not universal ones; run the test on your own incident history.

FunctionOutput classWho signs todayError surfaces inUnwind cost after that delayGoverning regime
IT / engineeringConfig change, ticket resolutionReviewer + CI, per changeMinutesLow, rollbackInternal change management
Finance / accountingInvoice-to-PO match, journal entryController at close; auditor annuallyDays to weeks; fraud median 12 monthsLow in-period, high post-closeFinancial reporting controls, audit
Customer supportReply to a customerNobody per message; QA samplesSecondsHigh; the statement may bindConsumer protection; misrepresentation
Sales / marketingOutbound message, CRM recordRep owns the number, not the messageWeeks, at pipeline reviewMedium; relationship and data hygieneMarketing and privacy law
HR / recruitingScreen, rank, rejectRecruiter, but rejections are silentOften neverVery high; statutory and retroactiveEU AI Act Annex III; NYC Local Law 144
Legal / contractsFiling, clause, adviceAttorney of record, on every filingDays to weeks; opposing counsel or judgeVery high; sanctions, public and personalProfessional duty; court rules

Read the table by column, not by row. The "who signs today" column is the one that predicts deployment success; the "output class" column is the one every competing guide sorts on.

Finance and Accounting: The Close Is the Countersignature

Finance is the strongest starting function for AI agent business applications in most companies, and the reason has nothing to do with the tasks being easy. It is that finance is the only department outside engineering that already runs a scheduled, mandatory, documented review of everything it produced: the close, with a named human attesting to it.

Best for: invoice-to-purchase-order matching, expense-policy checks, reconciliation exceptions, vendor-master hygiene, close-checklist preparation.

What the signer sees: the controller signs the period, not the transaction. That means the countersignature is real but coarse. An agent that posts 4,000 matched invoices gets one signature covering all of them, which only works if the agent's exceptions are surfaced separately and the match rate is monitored between closes.

Detection latency: the honest number here is uncomfortable. The Association of Certified Fraud Examiners' Occupational Fraud 2026: A Report to the Nations, covering 2,402 cases, found the median fraud scheme lasted 12 months before detection, and that tips, not controls, were the most frequent detection method at 43 percent of cases. Median loss was $104,000 per case overall, falling to $40,000 where the scheme was caught within six months. Those figures describe deliberate concealment rather than agent error, but they set an upper bound on how much confidence to place in "finance will catch it." Finance catches things at the cadence of its calendar, and the calendar is the control.

Unwind cost: low inside the period, materially higher after the books close and the number has been reported. That asymmetry gives finance agents a natural safe window, and it is the window your Field 2 measurement has to fit inside.

What we would grant on day one: read access to the ERP, write access to a staging table only, and an exception queue that a human clears. Not posting rights. The published number is somebody's signed attestation; an agent should reach it through a person, and the exception queue is where the real cost of invoice automation lives.

Customer Support: Fastest Detection, Most Binding Output

Support is where most companies want to start and where the countersignature is weakest. Every other function reviews a sample; support reviews a sample after the customer has already read it. There is no pre-send signature on a support reply, which means the agent's output is the company's final word by default.

Best for: deflection of known-answer questions, ticket triage and routing, drafting replies for agent review, summarising history before a human takes over.

What the signer sees: in most support organizations, nothing before send. Quality assurance reads a small sample of conversations after the fact. That is a countersignature that arrives too late to be one.

Why the binding matters: in Moffatt v. Air Canada, decided by British Columbia's Civil Resolution Tribunal on 14 February 2024, a customer relied on an airline chatbot that described a bereavement-fare process the airline's own webpage contradicted. Air Canada argued it could not be liable for information provided by the chatbot. The tribunal's response is worth quoting because it disposes of the entire "the agent said it, not us" defence in two lines: "This is a remarkable submission", the member wrote, before adding that it "should be obvious to Air Canada that it is responsible for all the information on its website." The tribunal found negligent misrepresentation and ordered $812.02 in total, of which $650.88 was damages.

The money is trivial. The holding is not. A chatbot's statement was treated as the company's statement, and the company's argument that the tool was somehow a separate responsible entity was described as remarkable — meaning unserious. Every support agent deployment inherits that holding.

What we would grant on day one: draft-and-suggest, not send. Where auto-send is genuinely required for volume reasons, restrict it to an allow-listed set of intents with fixed, human-written answer templates, and log the intent classification alongside the reply so the sample review can be aimed at the classifier rather than the prose. We set out the escalation structure for this in our read, say, do ladder for customer service automation.

Sales and Marketing: Nobody Signs the Email

Sales is the function where the countersignature question produces the most surprised faces. Ask a VP of Sales who signs off on an outbound email and the answer is usually a pause, then "the rep, I suppose." But the rep owns a quota number, not a message. Nobody reviews the message.

Best for: lead research and enrichment, CRM hygiene, meeting preparation briefs, call-note summarisation, sequencing follow-ups a human sends.

What the signer sees: the pipeline review, typically weekly or fortnightly. That is the department's real countersignature, and it inspects aggregates, not artifacts. A rep whose numbers look normal is not asked what the agent wrote.

Detection latency: weeks, and biased. Errors that reduce reply rates are invisible because non-response is the base case. An outbound message with the wrong company name or an invented product claim does not generate a complaint; it generates silence that looks exactly like ordinary silence. This is the missing-detector problem in its purest form.

Unwind cost: medium and mostly reputational, with one sharp exception. Where an agent writes to the CRM, bad records compound: every downstream forecast, territory assignment and comp calculation reads from a system of record the agent has been editing without review. Reversal there is a data-quality project, not an undo.

Regime: marketing and privacy law rather than sector regulation: consent, suppression lists, accurate claims. Lower stakes than HR or legal, which is precisely why sales tolerates agents with the least oversight and accumulates the most silent drift.

What we would grant on day one: full read on the CRM, write confined to enrichment fields the agent owns exclusively, no write to opportunity stage or amount, and outbound restricted to sequences a human approved as templates. Track the reversal cost of sales process automation before widening.

HR and Recruiting: The Countersignature That Never Comes Back

HR is the function where the agent looks most useful and the countersignature is most thoroughly broken. Screening, ranking and rejecting candidates is high-volume, rule-shaped work that agents handle well. It is also the only category on this list where the error usually never surfaces, and where the regime attaches statutory liability regardless of intent.

Best for: interview scheduling, job-description drafting, internal policy Q&A, onboarding checklists. Deliberately not: screening, ranking, or rejecting.

What the signer sees: a recruiter signs the offer. Nobody signs the rejections, and rejections are the overwhelming majority of the output. A candidate who was wrongly filtered out is not told and cannot appeal.

Regime, in detail. The EU AI Act classifies employment AI as high-risk in Annex III point 4, covering systems "intended to be used for the recruitment or selection of natural persons, in particular to place targeted job advertisements, to analyse and filter job applications, and to evaluate candidates," and separately systems used "to make decisions affecting terms of work-related relationships, the promotion or termination of work-related contractual relationships" (Annex III, artificialintelligenceact.eu). In New York City, Local Law 144 of 2021 "prohibits employers and employment agencies from using an automated employment decision tool unless the tool has been subject to a bias audit within one year of the use of the tool, information about the bias audit is publicly available, and certain notices have been provided to employees or job candidates" (NYC Department of Consumer and Worker Protection); enforcement began 5 July 2023, and the notice must be given ten business days before use.

Two things follow. First, the audit and notice obligations attach to the tool, so buying an agent platform does not transfer them. The employer holds them. Second, both regimes assume a decision trail exists. If your agent ranks candidates and you cannot reconstruct why, you do not have a compliance gap so much as an absence of the artifact compliance is about. Our guide to EU AI Act obligations for deployers covers the deployer-side duties in more depth.

What we would grant on day one: scheduling and drafting only, with any ranking output visible to the recruiter as a suggestion carrying its reasons, never as an ordering that silently removes candidates from view.

Legal is the one function on this list where the countersignature already exists, is mandatory, is individual, and is enforced by an outside body. An attorney signs every filing. That is a strong control, and the record shows it is not sufficient, because the signer can sign without verifying.

Best for: first-pass contract review against a clause playbook, document classification, discovery triage, summarising a matter for a partner.

Detection and consequence: errors surface when opposing counsel or a judge reads the filing, so within days to weeks, and then they surface publicly, attached to a named individual. The scale of this is now measurable. Damien Charlotin's AI Hallucination Cases database, which "tracks legal decisions in cases where generative AI produced hallucinated content – typically fake citations," listed 1,814 identified cases as of its 30 July 2026 update, 1,252 of them from the United States, with Canada at 200 and the United Kingdom at 61.

That number is the strongest available evidence for the argument in this piece. Legal has the best countersignature structure of any business function: one licensed human, personally accountable, on every artifact. And it has produced more documented AI failures than any other function precisely because the signature is a formality unless the signer can check the work. A signature that cannot verify is not a control. It is a name on a mistake.

What we would grant on day one: retrieval and drafting against an internal, closed corpus, with every citation resolved to a document the firm holds, and an automated check that refuses to output a citation it cannot resolve. The failure mode here is not the model being wrong; it is the model being fluent about something that does not exist.

IT and Engineering: Where the Countersignature Was Automated First

This is the control case for the whole argument. McKinsey found IT among the two functions where agent use is most commonly reported, and IT is also the function whose countersignature was rebuilt for machine speed a decade before agents existed.

Consider what an engineering change already passes through: a diff another engineer reads, an automated test suite, a linter, a staged rollout, monitoring with alert thresholds, and a rollback path measured in minutes. None of that was built for AI. It was built because humans ship broken code, and it happens to be exactly the structure an agent's output needs: machine-readable proposal, machine-executed check, cheap reversal, fast detection.

Best for: service-desk triage and resolution, log analysis, configuration drift detection, dependency upgrades, runbook execution.

The transferable lesson: the reason engineering absorbs agents smoothly is not that engineers are better at AI. It is that engineering had already answered all four fields of the Countersignature Test, with a named reviewer, minutes to detect, a cheap unwind and an internal-only regime, before the question was asked. Every other department is being asked to answer those four fields for the first time, under time pressure, while a pilot is already running.

If you want a target state for another function, describe engineering's control set in that function's language. What is the diff? What is the test suite? What is the rollback? A finance team that can answer those three questions has built a countersignature; one that cannot has bought an agent.

Timeline showing how long an agent error takes to surface in each of six business functions

Detection windows by function, on a log-scaled timeline. The rightmost case has no natural detector at all.

The Accountability Gap: The Builder Is Not the Signer

There is a structural defect underneath all six functions, and it is the one the Hacker News comment above was describing. In most companies the person who builds the agent and the person who answers for its output are different people in different reporting lines, and nothing formally connects them.

The builder is usually in operations, IT, or a technically-inclined corner of the department. The signer is the function head, the controller, the recruiter, the attorney. The builder tunes prompts and connectors; the signer discovers the agent exists when something goes wrong. That is the accountability gap, and it is a governance defect rather than a technical one, which is why buying a better agent framework does not close it.

Three symptoms tell you the gap is open in your organization:

  1. The agent runs under a human's credential. If the agent authenticates as the person who built it, it inherits that person's access, and its actions are indistinguishable from theirs in every log. We have argued at length that every agent needs an owner, a scope and an expiry as a first-class identity; the practical test is whether you can answer "which agent did this" from your audit trail without asking a person.
  2. Nobody in the owning function can describe what the agent is allowed to touch. Ask the controller which systems the finance agent can write to. If the answer requires fetching the builder, the signature is uninformed.
  3. The agent has no expiry and no review date. Controls that never expire are controls nobody re-examines. A signature renewed annually is a real one; a signature given once at launch is a launch approval.
Builder and function head in separate reporting lines before the fix, and connected through a registered agent identity after

The accountability gap before and after. The fix is a registration change, not a reorganisation.

Closing the gap does not require a reorganisation. It requires that the agent be registered as an asset of the function that owns the output, with the function head named as owner, and that the builder hold a delegated build role rather than de facto ownership. That single change makes the countersignature meaningful, because the signer now has standing to refuse.

A Worked Example: One Application, Two Departments

Take one output class, a drafted email that goes to an external party, and run the Countersignature Test on it twice, once in support and once in sales. Same agent, same model, same connector. The test produces different answers, and therefore different deployments.

FieldSupport: reply to an open ticketSales: outbound follow-up
Signer todayNone per message; QA samples after sendNone; rep owns the number, not the text
Surface timeSeconds; the customer repliesWeeks; pipeline review, and only in aggregate
Unwind costHigh: the statement may be a representation the company is held toMedium: reputational, plus CRM data drift
RegimeConsumer protection; negligent misrepresentationMarketing and privacy rules
Missing detector?No; the customer is the detectorYes; silence is the base case

The two deployments that follow are not variations of the same design. In support, the control has to be pre-send, because detection is fast but reversal is not available at all: draft-and-suggest, or auto-send confined to allow-listed intents with human-written templates. In sales, the control has to be a detector, because reversal is affordable but nothing surfaces the error: sample outbound messages on a schedule the way support samples tickets, and treat unusually low reply rates as a signal to read the text rather than to change the subject line.

Notice what did not determine the design: the agent's autonomy level, its model, or how many systems it touched. Those were identical. The department's existing accountability structure determined everything.

Now change one variable. Move the same drafted-email application into HR, as a candidate rejection. Signer: nobody. Surface time: never. Unwind cost: statutory and retroactive. Regime: Annex III and Local Law 144. The correct deployment is not a stricter version of the sales one. It is a different application entirely, or none.

What This Test Is Not

The Countersignature Test does not replace agent-side analysis, and reading it as a competitor to autonomy or blast-radius scoring will produce a worse deployment than either. The two operate on different objects.

Agent-side tests rank the agent: how reversible its actions are, how much it can reach, whether it acts without asking. We have published that ranking in detail, both as a four-test delegation ceiling for business agents and as a blast-radius ordering of individual use cases. Those tests are necessary. They tell you what the software is permitted to do.

The Countersignature Test ranks the function: whether the department that owns the outcome can detect, verify and reverse what the agent produced. It tells you whether permission is meaningful. An agent scored as low-blast-radius in a department that cannot detect its errors is not low-risk; it is unmeasured.

Run both. Where they agree, proceed. Where they disagree, whether that is a conservative agent in a department with no detector or an aggressive agent in a department with a machine-speed check, the department's answer should win, because the department is what remains after the pilot ends.

It is also not a maturity model. There is no ordering in which every company should adopt these functions. A company whose support organization already reviews every reply before send has a stronger support countersignature than most companies' finance close, and should start there.

When No Countersignature Can Be Built

Sometimes the honest output of the test is that this application should not run. Four patterns where we would not deploy, whatever the agent's capabilities:

The output has no detector and no signer. Silent, unreviewable output at volume. Candidate rejections are the canonical case. Adding an agent here increases throughput of a process nobody can audit, which is the opposite of the stated benefit.

The unwind is legally unavailable. Some outputs cannot be retracted in any meaningful sense: a statement to a regulator, a disclosure to a customer, a filing. The correct control is pre-output, and if pre-output review is too slow to be worth automating around, the automation has no case.

The signature would be uninformed. If the signer cannot verify the work in less time than doing it themselves would take, the countersignature is theatre. This is the legal-profession failure in the Charlotin data, and it is available to every function that treats approval as a click.

The regime forbids it in your jurisdiction and you have not read the regime. Employment, credit and essential-services decisions carry specific obligations, and the EU AI Act's Annex III also covers systems used "to evaluate the creditworthiness of natural persons or establish their credit score." Reading the rule is a prerequisite, not a follow-up.

There is a fifth case worth naming because it is common and undramatic: the process is broken and the agent will make it broken faster. If a department cannot describe the correct output in writing, an agent cannot be held to it, and the pilot will produce a disagreement about requirements rather than a result.

Building a Countersignature in 60 Days

Where the test returns a gap, the gap is buildable. This is the sequence we would run, and it is deliberately shorter than a governance programme.

Days 1–10. Name the output class and the owner. One class, not a department. The function head signs a one-page charter: what the agent produces, who owns it, what it may touch, when the grant expires. If nobody will sign the charter, that is the finding.

Days 11–20. Instrument before you automate. Log the current human process for the same output class: volume, error rate, how errors were found, how long they took to find. This is your Field 2 baseline and the only honest comparison you will get. Skipping it is why so many pilots report improvements nobody can defend.

Days 21–35. Run the agent in suggest mode. No writes to systems of record. The agent proposes; humans dispose; both are logged. Measure agreement rate and, more importantly, the character of disagreements. A 95 percent agreement rate with five percent catastrophic disagreements is worse than 80 percent with trivial ones.

Days 36–50. Build the detector, then narrow the grant. Whatever Field 2 said was missing, build it now: a sampling schedule, an exception queue, a reconciliation check, an alert threshold. Then scope the agent's access down to what the measured work actually required, which is almost always less than what was requested.

Days 51–60. Grant, with an expiry. Promote to limited autonomy inside the narrowed scope, with a review date, a named owner, and a kill path the owning function can trigger without filing a ticket with engineering. The design of approval gates as a control rather than a speed bump is what makes this stage hold.

The sequence assumes one application. Running three at once across three departments is how organizations end up with the accountability gap institutionalised, because the builder becomes the only person who understands all three.

Where LeapForce Fits

The Countersignature Test is a management exercise, and most of it happens in a room, not in software. What software has to supply is the evidence the signer needs and the boundary the charter promised. That is the layer LeapForce builds. Our platform treats every agent as a non-human identity with an owner, a scope and an expiry through Access & Identity, keeps the tamper-evident record of what an agent did and what it was refused through Observability & Audit, and chains human approval gates into event-triggered runs through Workflows. Our gateway rollout follows a deliberately unglamorous order: observe first, enforce second, optimize third. A department that starts by enforcing rules it has not yet measured usually enforces the wrong ones.

The honest limit: LeapForce does not decide who should sign in your finance close or your recruiting process, and it does not supply the departmental judgement this article is mostly about. Per-capability build status on our platform is disclosed as live, in development, or roadmap rather than presented uniformly, and readers evaluating us should ask for that breakdown directly.

Limits of This Analysis

Several things in this piece are weaker than they look, and it is better to say so than to let a reader find out later.

We have not run a controlled comparison. No LeapForce team member ran a side-by-side deployment of the same agent in two departments and measured the outcome difference; the framework is derived from published evidence, the failure record in the sources cited, and the structure of the accountability regimes themselves. Treat the four fields as a way to organise your own measurement, not as a validated predictor.

The detection-latency figures are class-typical, not universal. Twelve months is the ACFE median for occupational fraud, which is deliberate concealment and a harder detection problem than an agent's honest error. We use it as an upper bound on finance's detection confidence, not as an estimate of how long an agent mistake would hide.

Two sources a reader would expect are absent. The American Bar Association's Formal Opinion 512 on generative AI is the natural citation for the legal-profession duty of verification, and both it and the CanLII case record sat behind bot protection we could not pass, so the legal section rests on the Charlotin database and the Civil Resolution Tribunal's own published decision instead. We would rather name the gap than cite a summary of a document we did not read.

Jurisdiction coverage is uneven. The HR regime discussion covers the EU and New York City in detail because those texts were verifiable; Illinois, Colorado and several other jurisdictions have enacted or amended AI-in-employment rules on their own timelines and are not analysed here.

Where the framework may be wrong: if agent reliability improves faster than we expect, the value of a human countersignature falls and the binding constraint moves back to the agent-side tests. The strongest argument against this piece is that we are describing a two-to-three-year window rather than a permanent structure. We think the regimes outlast the models, but that is a judgement, not a finding.

 FAQ

Frequently asked questions

AI agent business applications are named, scoped software workers that carry a multi-step task across several systems without a human re-triggering each step, plus the accountability structure of the department that owns the output. Both halves ship together. The capability half decides what the agent can do; the accountability half decides what happens when it is wrong, and it is the half most buying decisions ignore.

Start where the department's existing sign-off already runs at machine speed and reversal stays cheap for the whole detection window. In most companies that is IT service desk work, then finance reconciliation and matching. McKinsey's November 2025 State of AI survey found agent use most commonly reported in IT and knowledge management, which fits: engineering had already automated its own review, testing and rollback long before agents arrived.

The company, and specifically the function that owns the output. In Moffatt v. Air Canada (2024 BCCRT 149), the airline argued it was not liable for what its chatbot said; the tribunal called that "a remarkable submission" and held that the company is responsible for all the information on its website regardless of whether it came from a static page or a chatbot. Internally, accountability should be assigned to the function head who signs that class of output, not to whoever built the agent.

Autonomy and blast-radius scoring rank the agent: how reversible its actions are, what it can reach, whether it acts unattended. The Countersignature Test ranks the function: whether the owning department can detect, verify and reverse the output. Both are necessary. An agent with a small blast radius sitting in a department with no error detector is not low-risk, only unmeasured, and the department is what remains after the pilot ends.

The visible cost is platform seats plus model usage, which scales with run volume rather than headcount. The cost most budgets miss is the countersignature: the reviewer time, the sampling schedule and the exception queue that make the output verifiable. Budget for the detector as a line item. A pilot that is cheap because nobody reviews the output has not measured anything, and its savings figure will not survive the first incident.

Sixty days is realistic for one output class in one department when the sequence is instrument, suggest-mode, build the detector, then grant with an expiry. The long pole is almost never model quality; it is producing the Field 2 baseline: what the current human error rate is and how errors were historically found. Departments that skip that step ship faster and cannot defend the result afterwards.

Safe is the wrong frame; the question is whether the regime's required artifacts exist. Regulated processes generally demand a reconstructable decision trail, a named accountable person, and evidence of testing. An agent that cannot produce those does not create a compliance gap so much as remove the artifact compliance is about. Where the regime attaches strict liability — as employment decisions do under New York City's rules, buying a vendor's tool does not transfer the obligation.

A department can build one, and increasingly does. The risk is not technical competence but the accountability gap: the builder becomes the de facto owner while the function head carries the consequence. Whoever builds it, register the agent to the function that owns the output, name that function's head as owner, give the agent its own identity rather than a person's credential, and set an expiry date on the grant.

Score AI agent platforms for business on what they give the countersigner, not on model choice or connector count. Four questions: can the agent hold an identity separate from a human's; can you reconstruct after the fact why it acted; can an approval gate be placed on a specific action rather than the whole workflow; and can a business owner revoke access without an engineering ticket. Platforms differ far more on those than on capability.

Enough to refuse. At minimum: what the agent proposes to do, what it read to decide, which systems the action touches, and what reversing it would take. McKinsey's survey found explainability was among the most-reported AI risks but not among the most-mitigated. That gap is exactly the countersigner's problem, because approval without a reason to approve is a signature on a mistake rather than a control.

Employment decisions are the clearest. The EU AI Act's Annex III point 4 lists systems used to recruit, filter applications and evaluate candidates, and separately those affecting promotion or termination, as high-risk. Annex III also covers creditworthiness evaluation and eligibility for essential public services. New York City's Local Law 144 adds a bias-audit and candidate-notice requirement for automated employment decision tools, enforced since 5 July 2023. Treat these AI agent use cases as regulated deployments, not pilots.

If the agent runs under that person's credential, it either breaks at offboarding or keeps running with orphaned access, both bad, in different ways. This is why agent identity should be separate from human identity from day one, with an owner recorded at the function level. Ownership that survives a departure is the difference between an agent that is a company asset and one that was somebody's personal script all along.

Ready to Govern Your AI?

Talk to LeapForce — one controlled layer for every AI tool, connector, model, and agent.

Thirty minutes · No pitch deck

Ready to turn AI experiments into measurable ROI?

Bring one outcome you'd like AI to move. We'll help you scope a pilot you can actually measure — and tell you honestly if it's not worth doing yet.

Comments