AI Workforce Planning: Supervision Is the Real Limit

An AI workforce is the population of software agents a company runs alongside its people, each holding a bounded job, a named human owner, and a defined level o

An AI workforce is the population of software agents a company runs alongside its people, each holding a bounded job, a named human owner, and a defined level of autonomy. Its size is not set by how many agents you can build. It is set by how much human attention you can afford to spend supervising them.

That is the part almost every article on this subject leaves out, and it is the part that decides whether your programme survives contact with a real quarter. We think the useful planning unit for an AI workforce is not headcount, agent count, or licence count. It is oversight load: the minutes of genuine human attention each agent consumes per week. Add up the load, divide by the review capacity you actually have, and you get a ceiling. Every organisation has one. Very few have measured it.

The question is already live in practice. An engineer posting on Hacker News in April 2026 under the handle dsteel described running fourteen agents across several frameworks and finding that none of the tooling answered the questions that mattered: which agent has authority over which domain, what happens when two disagree, who owns the escalation path. "Those are organizational problems, not technical ones," they wrote. In June, on a thread about coding agents, jpgvm put the same constraint in one line: there is "no way to keep up with their output at human review pace."

The short answer: Size your AI workforce by dividing your team's real weekly review capacity by the oversight load of each agent class. For irreversible, customer-facing work that usually lands near one agent per supervisor, not the dozens the category page implies.

Last updated: July 30, 2026.

Three supervision classes with weekly oversight load and how many agents one supervisor can carry

Oversight load per agent class, against one supervisor with six real review hours a week. Illustrative model.

One line of housekeeping before the argument. Nobody at LeapForce has run a controlled study of supervision ratios inside a customer's operations team, and this article does not pretend otherwise. The arithmetic below is a transparent model built from stated assumptions you can replace with your own numbers, and every external figure is attributed to a named source. Where the evidence runs out, the article says so rather than filling the gap with a confident sentence.

What an AI workforce actually is

An AI workforce is a managed population of software agents that hold bounded jobs inside a company: each one has a task it repeats, systems it may touch, an owner who answers for it, and a rule about how much it may do before a person looks. The term is doing real work only when all four of those are true. If any is missing, what you have is a tool with a personality, not a workforce member. Vendors also sell the same idea as digital labour, digital workers or an agentic workforce; the words differ, the planning problem does not.

Three modes get lumped together under the label, and they carry very different supervision costs.

ModeWhat it doesHuman involvementTypical oversight cost
AssistiveDrafts, summarises, retrieves; classic task automation with a language model attachedEvery output passes a person before it leavesLow per run, high in aggregate
AdvisoryAnalyses and recommends; the human decidesThe decision stays human, the reasoning is machineModerate, concentrated on hard cases
ExecutiveRuns a multi-step workflow and acts in live systemsHuman sets policy and reviews exceptionsHighest, and it scales with blast radius

The third mode is the one that changed. AI agents in the workplace have been drafting and summarising for three years. What is newly available is software that starts work without being asked, chains several steps together, and writes into systems of record. The 2026 AI Index Report from Stanford HAI reports that agent success on OSWorld, a benchmark of real computer tasks, climbed from roughly 12% to about 66% over the course of 2025, while noting that agents still fail around one attempt in three on structured benchmarks. On a harder benchmark built from realistic company tasks, TheAgentCompany, the most competitive agent completed 30% of tasks autonomously.

Read those two numbers together and you have the whole planning problem. A one-in-three failure rate is remarkable engineering and an unremarkable colleague. It is good enough that delegating work is rational. It is bad enough that nobody sane lets the output go unchecked. The gap between those two facts is filled by human attention, and human attention is the scarce input.

What an AI workforce is not

It is not a headcount substitution plan. It is not robotic process automation with better marketing. RPA replays a fixed script and fails loudly when the screen moves, while an agent improvises and fails quietly in prose that reads correct. It is not a multi-agent swarm of interchangeable general assistants; the useful ones are narrow, and orchestration frameworks solve message passing rather than authority. And it is not a technology programme. As dsteel put it on Hacker News, the unanswered questions are organisational: authority, disagreement, escalation.

Why headcount is the wrong first question

The first question most leadership teams ask about an AI workforce is how many people it replaces. That question cannot be answered before a prior one: how many agents can this organisation actually supervise at the quality bar its customers already expect? Until you have that number, every business case is arithmetic on top of an unmeasured constraint.

The evidence for the constraint keeps arriving from the operational end rather than the strategic one. Microsoft's 2026 Work Trend Index, which surveyed 20,000 full-time workers across ten countries between February and April 2026, reports that "the number of active agents in the Microsoft 365 ecosystem has grown 15x year over year, rising to 18x in large enterprises." The same research finds that organisational factors, meaning culture, manager support and talent practices, account for 67% of the reported impact of AI against 32% for individual factors. Capability is not the bottleneck. The surrounding system is.

That maps onto what practitioners describe. A developer posting to Hacker News in March 2026 under the handle gentle\_bubble wrote that AI coding tools doubled their team's pull-request volume while "turning senior engineers into full-time PR reviewers." The output went up. The supervisory capacity did not. The result was not more throughput; it was a queue.

This is why we treat AI workforce planning as a capacity-planning exercise rather than an automation-selection exercise. You are not choosing tools. You are allocating a fixed pool of expert attention across a growing population of things that need watching, and deciding, deliberately and in writing, which of them you are going to stop watching.

The supervision ratio, defined

The supervision ratio is the number of agents one human can oversee at an acceptable error rate, for a given class of work. It is not one number for a company. It is one number per class of work, and it moves when you change the design of the work, not when you change the model.

The question is older than the current wave. In Technology in Society in February 2022, Elissa Farrow published a study on determining the human to AI workforce ratio, running 110 participants from 36 organisations through structured futures workshops across five operating models, from a fully human workforce to an AI-led one with no humans. The output was a planning approach, not a measured ratio, which is itself informative: four years ago the field could already see that the ratio was the strategic variable, and it still has no empirical benchmark for it.

So the honest position is that no published number tells you how many agents a person can supervise. What you can do is compute your own, from four quantities you can observe this week:

SymbolQuantityHow to get it
RRuns per week for this agentCount them, or estimate from the task's current human volume
fFraction of runs a human must inspectSet by policy and reversibility, not by preference
mMean minutes per inspectionTime five real reviews with a stopwatch
SSupervisor's real weekly review capacity, in minutesTheir available hours minus their existing job

Oversight load per agent is L = R × f × m. The supervision ratio for that class is S ÷ L. That is the whole framework, and its usefulness is entirely in the discipline of measuring S honestly.

Most plans fail on S. A team lead nominated as the owner of three agents does not have forty hours of review capacity; they have their own job plus whatever is left. Six hours a week, or 360 minutes, is a defensible planning figure for someone who still has other responsibilities, and it is the figure used throughout the worked example below. If your supervisors genuinely have more, substitute it. If they have less, the ceiling falls proportionally and you should know that before you commit to a rollout date.

Oversight load: the arithmetic that sets your ceiling

Oversight load is the minutes of human attention an agent consumes per week, and it is what converts an abstract AI workforce plan into a staffing number. Three classes of work carry very different loads, because reversibility drives the review fraction more than accuracy does.

Take a mid-market operations team. One supervisor, 360 minutes of real weekly review capacity, three candidate agents. Every number in the table below is an assumption we are stating openly so you can replace it; none of it is measured field data, and quoting the right-hand column as a benchmark would be a misuse of it.

Supervision classExample workRuns/week (R)Reviewed (f)Min each (m)Load (L)Agents per supervisor
Read-only draftingMeeting summaries, research notes10010%330 min12
Reversible writeInternal CRM and ticket updates20015%4120 min3
Irreversible or customer-facingSent email, refund, published copy60100%5300 min1.2

The spread is the finding. The same supervisor can carry a dozen drafting agents or barely one agent that sends things to customers. Any plan that quotes a single company-wide ratio has averaged across that ten-fold difference and will be wrong in both directions at once: over-supervising the harmless agents, under-supervising the dangerous one.

Now build a portfolio. Suppose the team wants twenty agents live by the end of the year: twelve read-only, six reversible, two irreversible.

ClassAgentsLoad eachTotal load
Read-only drafting1230 min360 min
Reversible write6120 min720 min
Irreversible or customer-facing2300 min600 min
Total201,680 min = 28 hours/week

Twenty agents consume twenty-eight hours of expert review every week. At six hours per supervisor that is 4.7 people's supervisory capacity, permanently, for as long as the agents run. Nobody budgets for that, because the business case was written about the work the agents do, not the work they create.

Two things follow immediately, and both are more useful than a ratio.

First, the two irreversible agents, 10% of the population, consume 36% of the load. Blast radius, not agent count, drives the staffing model. This is the same conclusion we reached from the delegation side in our earlier analysis of the delegation ceiling, approached from the opposite end: that piece asks which work you may hand over, this one asks how many hand-overs you can afford to watch.

Second, if 28 hours a week is unaffordable, and in a twenty-person team it usually is, you have exactly three options. Run fewer agents. Hire supervisors. Or reduce the load per agent. Only the third one scales.

Raising the ratio without hiring supervisors

The supervision ratio is not a fact about your people. It is a design parameter, and two levers move it. Both work by changing the arithmetic rather than the staffing, which is why they are the only part of AI workforce planning that compounds.

Lever one: cut the reviewed share (f). Every review that exists because a human might catch a bad action is a review that a constraint could have prevented. If an agent must never issue a refund above a threshold, express that as a hard limit the agent cannot exceed rather than as a thing the reviewer watches for. On the reversible-write class above, moving f from 15% to 3% by putting three such constraints in policy takes the load from 120 minutes to 24, and the ratio from 3 agents per supervisor to 15.

The general rule is that you should forbid what you can forbid and review only what genuinely requires judgment. Detection is a running cost that scales with volume; prohibition is a one-time cost that does not. We have written the security-side version of this argument in our analysis of AI agent guardrails; the workforce-planning consequence is that the same constraint pays twice, once in risk and once in staffing.

Lever two: cut the minutes per review (m). Most review time is not judgment. It is reconstruction. Working out what the agent saw, which tool it called, what it changed, and why. If the action record answers those questions before the reviewer asks, a five-minute reconstruction becomes a two-minute confirmation. On the irreversible class, that takes the load from 300 minutes to 120 and the ratio from 1.2 agents to 3.

Apply both levers to the twenty-agent portfolio:

ClassAgentsLoad beforeLoad afterChange
Read-only drafting12360 min360 minunchanged
Reversible write6720 min144 minf: 15% to 3%
Irreversible or customer-facing2600 min240 minm: 5 min to 2 min
Total201,680 min744 min28 h to 12.4 h/week

The same twenty agents, the same people, the same risk appetite, and the weekly supervisory bill falls from roughly 4.7 people to about 2.1. The savings did not come from a better model. They came from writing down what is forbidden and instrumenting what happened.

That is the sentence we would put on the whiteboard. In an AI workforce, capacity is bought with policy and evidence, not with headcount.

What the labour-market data actually shows

The honest labour-market picture in mid-2026 is narrower and sharper than either the replacement story or the reassurance story. Displacement is real, measurable, and concentrated at the entry level, which is exactly where the future supervisors were supposed to come from.

The most careful evidence is Canaries in the Coal Mine?, published on 13 November 2025 by Erik Brynjolfsson, Bharat Chandar and Ruyu Chen of Stanford using high-frequency administrative payroll data from ADP. Their abstract states that "early-career workers (ages 22-25) in AI-exposed occupations experienced 16% relative employment declines, controlling for firm-level shocks, while employment for experienced workers remained stable." Two further findings matter for planning: adjustment happens through employment rather than pay, and the declines concentrate in occupations where AI automates rather than augments.

Set that beside the adoption base rate. A Federal Reserve FEDS Note published on 3 April 2026 puts firm-level AI adoption at roughly 18% of US firms as of December 2025, with the highest rates in professional services (around 33%) and financial services (around 30%), and single digits in accommodation and food services. Adoption is not general. It is concentrated in cognitive, analytical work. The same work entry-level knowledge staff used to do.

SourceFindingDate
Stanford Digital Economy Lab (ADP data)16% relative employment decline, ages 22-25, AI-exposed occupationsNov 2025
Federal Reserve FEDS Note~18% of US firms using AI; ~33% professional services, ~30% financeApr 2026
Stanford HAI AI Index88% organisational adoption; agent deployment still single-digit in many functions2026
Anthropic Economic Index"Over a third expect AI to be able to do most or nearly all of their work tasks next year"Jun 2026

The Anthropic Economic Index report of 26 June 2026 adds a useful corrective to the displacement story: in its data, higher-wage occupations show more turns per conversation and more output per turn, a pattern the authors describe as labour-augmenting rather than labour-displacing. Both things are true at once. Senior work is being amplified; junior work is being absorbed.

For an AI workforce plan, that combination has an uncomfortable implication that almost nobody states. The supervisory capacity your plan depends on is produced by a career pipeline that the same technology is thinning. If you cut the analyst intake because agents draft the analyses, you have also cut the supply of people who will be qualified to review agent output in five years. That is not a moral argument. It is a capacity argument, and it belongs in the workforce plan next to the licence costs. It is also the one place where upskilling budgets earn their keep for a reason nobody puts in the deck: reskilling mid-career staff into supervisory roles is how you refill a pipeline that job displacement at the entry level has drained.

Why supervision decays as the population grows

Supervision does not scale linearly, and this is the best-evidenced part of the whole subject. It just comes from a literature most AI strategy decks have never read. Human factors research has been measuring what happens to people asked to watch reliable automation for more than forty years, and the results are consistently unflattering.

Lisanne Bainbridge's Ironies of Automation, published in Automatica in 1983, is the origin text. Reviewing vigilance research, she concluded that "it is impossible for even a highly motivated human being to maintain effective visual attention" toward a source where very little happens, for more than about half an hour. Her wider argument is the one that should worry anyone drawing an org chart with agents on it: automation removes the routine work that kept the operator's skill and situational awareness current, then relies on that same operator to intervene correctly when something rare goes wrong.

The effect has a name and a mechanism. Raja Parasuraman and Dietrich Manzey's 2010 integrative review in Human Factors (volume 52, issue 3) concluded that automation complacency "occurs under conditions of multiple-task load, when manual tasks compete with the automated task for the operator's attention," is found in novices and experts alike, and "cannot be overcome with simple practice." Every word of that describes the supervisor in your AI workforce plan: a person with their own job, watching several agents, whose attention is exactly the resource being competed for.

Three consequences for planning follow, and they are not solved by training:

  1. Reliability makes supervision worse, not better. An agent that is right 97% of the time trains its reviewer to approve. The failure mode of a good agent is a rubber-stamp reviewer.
  2. Adding agents to a supervisor degrades all of them. Multiple-task load is the specific condition under which complacency appears. Two agents per person is not half the risk of four; it is a different regime.
  3. Skill fade is a scheduled event. A reviewer who has not done the underlying work for a year reviews worse than one who did it last month. Rotation is a control, not a perk.

This is why we treat the review fraction f as a policy variable rather than an aspiration. Telling a supervisor to check everything produces an unchecked system with a signature on it. Telling them to check the 3% that policy cannot constrain produces a reviewed system.

The strongest objection: maybe human review is the wrong gate

The sharpest argument against everything above is that the whole supervision-ratio framing preserves a control that has already stopped working. It deserves a hearing, because it is being made seriously.

A June 2026 arXiv preprint titled The End of Code Review: Coding Agents Supersede Human Inspection argues that "the naive integration in which agents write code and humans remain the mandatory reviewers is a dead end because it neither provides meaningful assurance nor scales with AI-assisted throughput." The authors' claim is that every stated goal of review can be met by agents at lower cost and higher throughput. It is a preprint, it is about software specifically, and it is contested. But the underlying observation matches both the Hacker News voices quoted earlier and the vigilance literature. A rubber-stamp review is worse than no review, because it manufactures false assurance and a signature.

We think the objection is right about the diagnosis and incomplete about the remedy. Two things are being conflated: inspection, which is a human reading output, and assurance, which is evidence that the action was permitted, bounded and recorded. Inspection genuinely does not scale. Assurance does, because it is machine-generated. Human-in-the-loop is not one setting; it is a dial with a cost attached at every position. The correct response is not to keep humans reading more output; it is to move as much of the burden as possible from f (how often a person looks) into constraints and records, and to reserve human judgment for the cases where the organisation would not accept a machine's answer regardless of accuracy: irreversible spend, customer harm, legal exposure, anything you would have to explain to a regulator.

Note what that does to the arithmetic: it is lever one and lever two, argued from the opposite direction. The disagreement is about whether the residual human gate goes to zero. On current evidence, and given that the EU AI Act obliges deployers of high-risk systems to assign oversight to competent humans, it does not.

The workforce register: four fields that make the ratio computable

You cannot plan a workforce you have not written down. Most companies already keep something for individual agents: an owner, a purpose, a set of permissions. We have argued elsewhere for that per-agent record in detail, in our earlier analysis of the AI employee personnel file. What almost nobody keeps is the population-level view: a register whose rows sum to a staffing number.

The difference matters. A personnel file answers "what is this agent allowed to do?" A workforce register answers "can we afford to run all of them?" Four fields turn the first into the second:

FieldWhy the plan breaks without it
Supervision classReversibility, not accuracy, sets the review fraction. Without a class, every agent gets averaged treatment.
Reviewed fraction (f)The number you are actually deciding. Left implicit, it defaults to "whatever the reviewer has time for".
Mean review minutes (m)Measured with a stopwatch, not estimated. This is where instrumentation pays back.
Named supervisor and their declared capacity (S)An owner without a stated weekly capacity is a name on a form, not a control.

A worked register row, filled in, for the customer-refund agent from the example:

FieldValue
AgentRefund processor, tier-1 support
Supervision classIrreversible, customer-facing
Runs per week60
Reviewed fraction100%, reducing to 20% once the value cap and duplicate-refund constraint are enforced in policy
Mean review minutes5.0 at launch; target 2.0 once the action record includes the ticket, the prior refunds and the policy decision
SupervisorNamed support lead, 6 hours declared weekly review capacity
Load at launch300 min/week — this agent alone consumes 83% of one supervisor
Load at target24 min/week (60 runs x 20% reviewed x 2 minutes)
Exit planOwnership transfers to the support lead's successor; credentials are the agent's own and are revoked, not inherited

That last row is not decoration. When the owner leaves, an agent whose credentials belong to a departing person becomes either an orphan or a security incident. Cloud Security Alliance research published on 20 May 2026 on non-human identity and agentic AI governance reports that 51% of organisations have no clear ownership of AI identities and only 20% have formal processes for offboarding and revoking API keys.

Counting the agents you already run

Before you plan an AI workforce, count the one you have. Most companies of any size are already running agents that no register knows about, and the count is not a footnote. It is the denominator of every ratio in this article.

The same Cloud Security Alliance research reports that non-human identities, meaning service accounts, API keys, OAuth tokens and agent credentials, outnumber human users by an average of 45 to 1, reaching 144 to 1 in some cloud environments, and that more than 16% of organisations do not track the creation of AI-related identities at all. Those ratios are not directly comparable to an agent headcount; most non-human identities are ordinary service accounts. But they establish the shape of the problem: the population is large, growing, and largely unowned.

A workable discovery pass, in the order that produces answers fastest:

  1. Follow the money. Pull every SaaS and model-provider charge on corporate cards and expense claims for the last two quarters. Personal-plan subscriptions expensed as software are the clearest signal of ungoverned automation.
  2. Follow the credentials. Enumerate API keys, service accounts and OAuth grants in your identity provider and major SaaS platforms, and sort by last-created rather than last-used. New non-human identities with no ticket behind them are agents nobody registered.
  3. Follow the traffic. Egress logs to model-provider endpoints tell you which machines are calling models and how often, whether or not anyone declared it.
  4. Ask, without penalty. A two-week amnesty window that asks teams to declare what they run finds things the first three passes miss, and only works if declaring costs nothing.
  5. Classify what you found. Every discovered agent gets a supervision class and a named owner within a week, or it gets switched off. Undecided is a decision.

We have covered the ungoverned end of this in more depth in our analysis of shadow AI. For workforce planning the point is narrower: an unregistered agent still consumes supervision, it just consumes it unpredictably and after something goes wrong.

A 90-day sequence for sizing your first cohort

Ninety days is enough to produce a defensible AI workforce plan for one function, and not enough to roll one out across a company. The sequence below deliberately spends the first month measuring rather than building, because every downstream number depends on S and m being real.

Days 1-30: measure the denominator.

  • Pick one function with high-volume, documented, repeatable work.
  • Run the discovery pass above and produce a count of what already runs there.
  • For every candidate task, record R from the current human volume.
  • Time five real reviews of a similar artifact with a stopwatch to get an honest m. Do not estimate it. An estimate carries an unknown error into every later step of the plan, and m appears in every load figure you will quote to a budget holder.
  • Ask each nominated supervisor to declare S in writing, net of their existing job. Write the number down where their manager can see it.

Days 31-60: onboard one agent per class, not one per idea.

  • Ship one read-only agent and one reversible-write agent. Leave irreversible work alone this quarter.
  • Calibrate before you supervise. For a low-volume agent, review every run for two weeks. For anything above roughly 25 runs a week, review a fixed sample. Fifty consecutive runs is enough to see the recurring failure shapes, and it beats committing to a 100% pass that quietly becomes a second full-time job.
  • Record every correction. The pattern in the corrections is your policy backlog: each recurring correction is a candidate constraint that removes a review permanently.
  • After two weeks, drop f to the level the correction data supports and keep recording.

Days 61-90: convert corrections into constraints, then publish the ratio.

  • Implement the top five constraints from the correction log. Re-measure f.
  • Instrument the action record until a reviewer can answer what happened, why, and what changed without opening another system. Re-measure m.
  • Publish the register: every agent, class, f, m, load, owner, declared capacity. One page.
  • Decide the next cohort by remaining capacity, not by enthusiasm. If the register says 40 minutes a week are left, that is one more reversible agent, not four.

The order is the point. Measuring first is what stops the programme becoming what gentle\_bubble described: doubled output, unchanged review capacity, and senior people converted into full-time reviewers by accident.

What an AI workforce actually costs

The licence is the small number. In a first-year AI workforce budget, the two largest lines are usually the supervisory time nobody costed and the integration work nobody scoped. The first of those is computable from the load model above.

Cost lineHow to size itFrequently missed because
Model and platform spendRuns per week times cost per run, per agentIt is the only line vendors quote
Supervisory timeTotal weekly load in hours times loaded hourly cost, times 52It is spent by people who already have jobs
Integration and connectorsPer system of record the agents write intoScoped as "an API call" until permissions arrive
InstrumentationBuilding the action record that cuts mTreated as nice-to-have; it is the payback lever
Policy and legal reviewPer supervision class, not per agentDeferred until a customer complains
Rework from unreviewed errorsError rate times cost per error times unreviewed volumeNever modelled at all

Put a number on the supervisory line for the twenty-agent portfolio. At 28 hours a week and an assumed loaded cost of $60 an hour, and substitute your own, that is $1,680 a week, or roughly $87,000 a year, of expert attention. After both levers, at 12.4 hours a week, it is about $38,700. The delta, near $49,000 a year for one function, is what the constraints and the instrumentation are worth, and it is a far more honest ROI case than any productivity multiplier, because it is arithmetic on your own measured inputs. It is only half a business case, and deliberately so: the benefit side has to come from the function's own throughput and quality data, which nobody outside your company can supply and no vendor should be trusted to estimate for you.

We have taken the wider version of this apart in our analysis of enterprise AI implementation cost. The workforce-specific addition is simply that supervision is a recurring operating cost with a headcount equivalent, and it should appear in the plan with a name against it.

What the law already assumes about your ratio

Regulation has already decided that human oversight of AI must be staffed by named, competent people. That makes the supervision ratio is not only an operating question but a compliance one. If you cannot say who oversees which agent and with what capacity, you cannot answer a regulator either.

Article 14 of the EU AI Act requires that high-risk AI systems be designed so "they can be effectively overseen by natural persons during the period in which they are in use", with measures "commensurate with the risks, level of autonomy and context of use". The people doing the overseeing must be able to properly understand the system's capacities and limitations, remain aware of automation bias, correctly interpret output, decide not to use the system or to disregard, override or reverse it, and intervene or stop it. Read that list against a reviewer carrying four agents on top of their day job.

Article 26 is the one that binds deployers rather than builders, and it is explicit: "Deployers shall assign human oversight to natural persons who have the necessary competence, training and authority, as well as the necessary support." It also requires deployers to keep automatically generated logs for at least six months, and, for workplace deployment, to inform workers' representatives and affected workers before the system is put into use.

InstrumentWhat it asks of an AI workforceStatus
EU AI Act Art. 14Systems built so natural persons can effectively oversee themBinding, phasing in through 2027
EU AI Act Art. 26Named overseers with competence, training, authority; logs kept at least six months; workers informedBinding, phasing in through 2027
NIST AI Risk Management FrameworkGovern, Map, Measure, Manage; inventory and accountability for AI in useVoluntary; RMF 1.0 published January 2023, Generative AI Profile July 2024
ISO/IEC 42001Auditable AI management system, including roles and controlsVoluntary certification

One caveat before you over-apply this: Articles 14 and 26 bind high-risk systems, and most internal drafting or ticket-triage agents are not high-risk under the Act's classification. The obligations bite when an agent touches employment decisions, credit, essential services or the other Annex III categories. Read the classification before you either panic or relax. The NIST AI Risk Management Framework is voluntary and useful here mostly for its first function: Govern, which is where accountability and roles live. The through-line across all four instruments is the same as the through-line of this article. An oversight arrangement that exists on paper but has no capacity behind it satisfies nobody, including you.

Where LeapForce fits

LeapForce does not sell an AI workforce, and we would be suspicious of anyone who says they do. What we build is the layer underneath it: one governed endpoint for every model and agent, so that identity, policy, routing, cost attribution and audit are enforced on the call rather than in a document. That is precisely the layer the two levers depend on. You cannot cut the reviewed share without a place to express constraints, and you cannot cut minutes-per-review without a complete record of what the agent did.

Concretely, and with our build status labelled honestly: gateway endpoints, tracing and SSO across AI surfaces are LIVE; vaulted provider keys, inline data-loss prevention and dollar budgets are IN DEV; shadow-AI discovery and compliance evidence packs are ROADMAP, which is why the discovery pass in this article is written to be run with the identity, finance and network tools you already have. Our rollout model on the AI Gateway is deliberately sequenced the same way as the 90-day plan above, observe first, enforce second, optimize third, because a constraint written before you have measured the traffic is a guess, and a guess in policy costs more supervision than it saves. Agent identity is treated as first-class throughout: every non-human actor gets an owner, a scope and an expiry, which is what makes a register more than a spreadsheet.

Where this is still uncertain

Several things in this article are models rather than measurements, and the difference matters if you are about to spend money.

  • The load arithmetic is a framework, not a benchmark. We have not run a controlled study of supervision ratios inside a customer's operations team, and we do not present the 12 / 3 / 1.2 figures as empirical. They are what the stated assumptions produce. Replace R, f, m and S with your own measured numbers and the structure holds while the outputs change.
  • No published human-agent ratio benchmark exists. Farrow's 2022 scenario research established the question; nothing since has answered it with field data. Anyone quoting an authoritative ratio is quoting a vendor.
  • The identity ratios disagree with each other. Published figures for non-human to human identities range from roughly 45:1 to 144:1 depending on methodology and environment. That spread is itself the finding: the population is not being counted consistently, so treat any single figure as an order of magnitude.
  • Agent benchmarks do not transfer to your work. OSWorld and TheAgentCompany measure general computer and company tasks. Your failure rate on your systems, with your data, is an empirical question only your own two-week calibration answers.
  • The labour-market evidence is recent and contested. The Stanford entry-level finding is careful and well identified, but it is one dataset over a short window, and reasonable economists read the same period differently.
  • Sources we could not verify. The US Census Bureau's Business Trends and Outlook Survey is the most-cited source for firm-level AI adoption, and census.gov refused every fetch we attempted, including a browser session. We have used the Federal Reserve's analysis of the same survey instead and flagged it rather than quoting figures we could not confirm at source. Similarly, the Parasuraman and Manzey review is cited by journal, volume and year without an outbound link, because the publisher blocks automated access; the abstract text quoted was verified through a scholarly index.
  • This article contains no first-hand test. Route taken honestly: no LeapForce measurement or engagement data is presented as evidence anywhere above.
 FAQ

Frequently asked questions

Yes, but the number depends entirely on what the agents do. Using a model of 360 minutes of real weekly review capacity, a supervisor can carry roughly a dozen read-only drafting agents, about three agents that make reversible internal writes, and barely more than one agent that takes irreversible or customer-facing action. Reversibility drives the review fraction more than accuracy does, so classify by blast radius before you count.

No. Robotic process automation replays a fixed script against a fixed interface and fails visibly when the screen changes. An AI workforce agent interprets, improvises and chains steps, which means it fails quietly and in fluent prose. The practical difference for planning is that RPA failures are detected by monitoring and agent failures are detected by reading, and reading is the expensive one.

Run four passes in parallel: card and expense charges for model providers and AI SaaS; API keys, service accounts and OAuth grants in your identity provider sorted by creation date; network egress to model endpoints; and a no-penalty amnesty window asking teams to declare what they run. Anything discovered gets a supervision class and a named owner within a week or gets switched off.

Whoever the register says, which is why the register needs a named successor before the departure rather than after. Cloud Security Alliance research from May 2026 found 51% of organisations have no clear ownership of AI identities and only 20% have formal processes for offboarding and revoking API keys. The structural fix is that an agent holds its own credentials with its own owner, scope and expiry, so that offboarding a person revokes their access without orphaning the agent.

For high-risk systems, yes, and it is specific about staffing. Article 26 requires deployers to "assign human oversight to natural persons who have the necessary competence, training and authority, as well as the necessary support", keep automatically generated logs for at least six months, and inform workers' representatives before a high-risk system is used at work. Article 14 requires the system itself to be designed so those people can effectively override, disregard or stop it.

Model spend is usually the smallest of the three big lines. For the twenty-agent portfolio modelled here, supervisory time alone runs about 28 hours a week, or roughly $87,000 a year at a $60 loaded hourly cost, before licences, integration or instrumentation. Applying constraint and record-keeping improvements cuts that to about 12.4 hours a week, near $38,700. Budget supervision as a recurring operating cost with a headcount equivalent, because that is what it is.

Entry-level knowledge work in AI-exposed occupations, on the current evidence. Stanford's Digital Economy Lab, using ADP payroll data through late 2025, found early-career workers aged 22 to 25 in AI-exposed occupations experienced 16% relative employment declines while experienced workers held steady, with effects concentrated where AI automates rather than augments. The planning implication is awkward: that is the same pipeline that produces future supervisors.

Three that are not on most training menus. First, the ability to specify a constraint precisely enough to be enforced rather than watched. Second, enough current hands-on familiarity with the underlying work to spot a plausible-but-wrong output. Human factors research going back to Bainbridge in 1983 shows monitoring degrades the skill needed to intervene. Third, the judgment to decide which failures are acceptable, which is a management decision rather than a technical one.

Do not lead with a productivity multiplier. Lead with the load model: here is the work, here is the volume, here is the review fraction and the measured minutes per review, here is the resulting supervisory cost, and here is what that cost becomes after we implement the constraints the calibration period identifies. The delta between those two numbers is defensible because every input is measured in your own organisation, and it survives the question a multiplier never does. Where does the money actually come from?

Usually because the pilot was staffed at a supervision level production cannot afford. During a pilot, an enthusiastic owner reviews everything and the results look excellent. In production, the same review fraction requires capacity nobody allocated, so either the reviews stop happening or the rollout does. Measuring supervisory capacity before the pilot, and setting the target review fraction as an explicit exit criterion, prevents both outcomes.

Build, but build the measurement first. The technology will keep moving, and nothing in the load model becomes obsolete when models improve. Better agents lower the review fraction, which is the input the framework already exposes. Waiting costs you the calibration data, which takes calendar time to collect and cannot be bought. Starting with one read-only agent and an honest stopwatch is a cheaper way to learn your ratio than a company-wide programme.

Ready to Govern Your AI?

Talk to LeapForce — one controlled layer for every AI tool, connector, model, and agent.

Thirty minutes · No pitch deck

Ready to turn AI experiments into measurable ROI?

Bring one outcome you'd like AI to move. We'll help you scope a pilot you can actually measure — and tell you honestly if it's not worth doing yet.

Comments