An AI employee is software that owns a recurring business function: it holds its own credentials, starts work without being asked, writes into your systems of record, and leaves a trail you can reconstruct. Miss one and you have a feature, not an employee.
Our position at LeapForce is that the category only becomes useful when you take the employment metaphor literally instead of using it as marketing. A real employee arrives with an identity, a job description, a scope of access, a supervisor, and a record. If you cannot produce those five things for the software you just bought, you did not hire anyone. That framing came out of watching where the practical arguments actually happen. On Hacker News in February 2026, a developer posting as Quothling described working in a NIS2-regulated field where, as they put it, "we can't give AI's any sort of real access because of the security risk", and then noted that the same company runs internal tools where an AI mistake would cost nothing but annoyance. One organisation, two opposite correct answers, and no framework in the vendor literature for telling them apart.
The short answer: Treat an AI employee as a hire with a personnel file — its own identity, one named function, least-privilege scope with an expiry, a named human supervisor, and an action log that outlives it — and classify each of its tasks by how long and how many people it takes to undo, not by how autonomous you feel comfortable being.
Last updated: July 30, 2026.
We have not run a controlled deployment of a commercial AI employee product and are not reporting one here; everything below is built from public regulation, published research, vendor documentation we fetched and dated, and practitioner accounts we link to directly. Reddit was unreachable behind its bot protection during research, so every field voice quoted here comes from Hacker News, which skews technical relative to the operations and finance teams who actually buy this software.
The Personnel File test: five records you must be able to produce on demand for any AI employee.
What an AI employee actually is
An AI employee is a software system assigned one recurring business function, given its own credentials and a bounded set of permissions, and judged on whether that function stops arriving in a human's queue. The working definition matters because the market is full of things wearing the label. Gartner called this "agent washing" in a June 2025 press release. It means rebranding assistants, robotic process automation and chatbots without substantial agentic capability. Gartner estimated that only about 130 of the thousands of agentic AI vendors on the market are real.
The unit of adoption is the giveaway. Automation tools are sold by the task: extract this field, move this row, send this template. An AI employee is sold by the function: invoice intake, tier-one triage, supplier onboarding, expense review. That difference changes who owns the outcome, what access the software needs, what happens when it is wrong, and, the part this article is about, what you have to prove afterwards.
Adoption is real but shallower than the noise suggests. In McKinsey's global survey published in November 2025, 23 percent of respondents said their organisation was scaling an agentic system somewhere in the enterprise and another 39 percent had begun experimenting. In no single business function did more than 10 percent report scaled agents. The honest read is many pilots, a thin layer of production, and a governance gap sitting exactly where the two meet.
Three capabilities separate an AI employee from automation you already own. It initiates: a schedule, an inbound event or a threshold starts the work, not a person typing. It decides: inputs vary and the software chooses a path rather than following a fixed branch. It writes: output lands in the CRM, the ledger, the ticket system or the customer's inbox, not in a chat window for someone to copy out. Remove the third and you have a very good assistant. Remove the first and you have a tool.
The Personnel File test: five records or it is not an employee
The Personnel File test is our answer to a category that keeps getting defined by vibes: an AI employee qualifies if, and only if, you can produce five records for it on demand. Every one is something you already hold for a human hire, and every one has a counterpart in your existing stack. If a vendor cannot tell you where each record lives, the product is a feature with a headcount metaphor stapled to it.
| Record | Human equivalent | What it is for an AI employee | How you check it |
|---|---|---|---|
| Identity | Employee number, staff account | A credential that belongs to this agent alone — not a shared service account, not a person's login | Ask who authenticated. The answer must be one name, not "the automation user" |
| Job description | Role definition, objectives | One recurring function with a defined trigger, defined inputs and one defined output | Write it in a sentence. If it needs three, it is three jobs |
| Scope | Building pass, system entitlements | Least privilege, granted on a start date, with an expiry and an action-level limit | List every system it can write to. Then remove one and see if anyone notices |
| Supervisor | Line manager | A named human owner, plus an approval gate on the actions that cannot be undone | The question "who allowed this agent?" has exactly one answer |
| Record | HR file, appraisal history | An action log that survives the agent's deletion and reconstructs who authorised what, on whose behalf | Pick a random action from last month and try to explain it end to end |
The five records are not equally hard. Identity and scope are solved problems borrowed from identity and access management. The job description is a management discipline rather than a technical one, and it is where most deployments actually fail. The supervisor record is cheap to create and expensive to skip. The action record is the one people postpone until a regulator, an auditor or a customer asks a question that cannot be answered.
The OWASP Non-Human Identity Top 10 for 2025 ranks improper offboarding as the number one risk to non-human identities, ahead of secret leakage and over-privileged accounts. The order is instructive. The community that spends its time on machine credentials thinks the most dangerous moment in a non-human identity's life is the moment it should have ended and did not. Almost every guide to this category covers the hiring; very few cover the leaving.
We have written about the identity half of this before, in our earlier analysis of non-human identity and why every agent needs an owner, a scope and an expiry. The personnel file extends that from a security control into a management artefact, because the people signing off on a digital workforce are usually not the people who run identity governance.
What an AI employee is not
An AI employee is not a chatbot, not robotic process automation, not a copilot, and not a person. Each of those four confusions costs money in a different way. Getting the negative definition right stops you buying the wrong thing and blaming the category. Some vendors sell the same thing as a digital worker rather than an AI employee; the labels are interchangeable, and the test below does not care which is on the invoice.
It is not a chatbot. A chatbot answers inside the conversation; the state of your business is unchanged when the window closes. The dividing line is write access, which we have argued in detail in our piece on where write access draws the line between AI agents and chatbots. A chatbot that drafts a refund email is a chatbot. The same system with the ability to issue the refund is a different risk class entirely, and it should have arrived through a different approval.
It is not RPA. Robotic process automation follows a recorded path and breaks when the path changes. An AI employee is supposed to absorb variance: a supplier who sends a PDF instead of a CSV, a customer who describes a problem in the wrong words. That tolerance is exactly why it needs supervision. Deterministic software fails loudly and identically; probabilistic software fails quietly and differently each time, which is why the record matters more here than it ever did for RPA.
It is not a copilot. A copilot sits next to a person and multiplies them; accountability stays with the human who pressed send. An AI employee removes the human from the loop for a defined set of actions, which moves accountability up to whoever authorised the scope. That is a governance change disguised as a productivity purchase.
It is not a person. On Hacker News in July 2026, the commenter JohnFen dismissed the whole category with "genAI is a tool, not a person." As a statement about moral and legal status that is simply correct: software cannot hold a duty of care, be disciplined, or be sued. Everything the employment metaphor buys you is operational: a checklist of records that already exists, and that most deployments skip. Everything it does not buy you is legal, and pretending otherwise is how organisations end up with an accountability hole where a manager should be.
Where the employee metaphor genuinely breaks
The employment framing has a serious technical objection against it, and an article that hides that objection is selling something. The counterargument says an agent is not a new class of identity at all; it is an application acting on a user's behalf, and it should sit inside the authorisation model you already run.
That case was made cleanly on Hacker News in October 2025 by shawneechase, who argued that "Agents are not a new class of identity" and that the right question when an agent acts is not who the agent is but which user is being acted for and what that user is allowed to do. The post names two anti-patterns worth writing on the wall: broad-permission service accounts, which turn every agent bug into a privilege-escalation event, and hardcoded user credentials, which hand the agent everything the human could do, destructive parts included.
We think both models are right, in different places, and the deciding factor is whose authority the action draws on:
| Question | Agent-as-app (delegated authority) | Agent-as-worker (own authority) |
|---|---|---|
| Who is the action taken for? | A specific human, in session | The company, on a schedule |
| Where do permissions come from? | The user's existing entitlements, scoped down | A grant made to the agent itself |
| What starts the work? | A person asks | An event, a timer or a threshold |
| Who is asked when it goes wrong? | The user who invoked it | The named owner of the agent |
| Typical example | A research assistant summarising a document you opened | Overnight invoice intake with nobody at a desk |
The failure mode is mixing them. An agent running on a schedule with nobody in session cannot borrow a user's authority, because there is no user. So teams paste in a human's long-lived token, and the audit trail now shows a person doing things at 3am. That one shortcut breaks the fifth record, and it is the most common cause of a log that cannot answer "on whose behalf".
The practitioner move is to give the agent its own footprint. A Hacker News commenter posting as rida described running a personal agent in a container and said what made it usable was "giving the agent its own identity. Separate email, separate 1Password vault" plus a read-only service account. That is a one-person version of what an enterprise identity team would build, arrived at independently because the alternative kept breaking.
Prerequisites: what has to exist before day one
Before you onboard an AI employee, five things need to exist, and none of them are the product. Teams that skip this list spend their pilot debugging their own organisation rather than the software, which is a large part of why Gartner expects more than 40 percent of agentic AI projects to be cancelled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls.
- A written definition of the function, with a measurable output. Not "help with support", but "acknowledge every inbound support email within 10 minutes with a correct category and a first-line answer, or escalate." If nobody can write the sentence, nobody can tell later whether it worked.
- A baseline you took before the agent existed. Volume per week, current handling time, current error rate, current cost. You cannot claim a saving against a number you never measured, and the review conversation in month three will otherwise be entirely vibes.
- A named owner with the authority to switch it off. For high-risk systems the EU AI Act requires exactly this, in the wording quoted in the legal section below. Outside high-risk use the principle still holds: an owner who cannot stop the thing is a nominee, not a supervisor.
- An identity and access model that can issue a credential to something that is not a person. Provisioning and deprovisioning a digital worker is role-based access control applied to an account nobody will ever log into. If your only options are a human account or a shared automation login, fix that before the agent arrives.
- A place the action log will land, with a retention period already decided. Deciding this after an incident is how organisations discover their logs rotated at 14 days.
The prerequisite people most often argue about is the second one. It feels like bureaucracy in week one and it is the only thing that makes month three honest.
Supervision classes: which work is safe to leave alone
Most buying guides ask whether you are comfortable with full autonomy. That is a taste question, and taste is not a control. The better question is mechanical: how long, and how many people, does it take to undo? Answer that per action and the supervision class falls out on its own.
| Class | Undo test | What it covers | Control required |
|---|---|---|---|
| Class 1 — run and report | Under an hour, one person, no external party involved | Drafting, summarising, tagging, classification, internal research, monitoring and alerting | Post-hoc log review; sampled human QA; no gate |
| Class 2 — draft and hold | Requires contacting someone outside the team or the company | Outbound customer email, public posts, CRM stage changes, ticket creation in a customer's system, scheduling with third parties | Named human approves before execution; approval recorded with the action |
| Class 3 — dual control | Cannot be undone, or undoing it is itself a reportable event | Payments and refunds above a threshold, contract execution, production configuration changes, deletion of records, anything affecting a person's employment | Two named humans; pre-authorised limits; separate audit review |
The classification is per action, not per agent. The same support agent can sit in Class 1 when it tags a ticket, Class 2 when it replies to the customer, and Class 3 if it issues a refund. That is precisely why "is this agent autonomous?" is the wrong unit of analysis. We have taken this apart in more depth in our analysis of setting a real autonomy budget for autonomous AI agents.
Two things push an action up a class regardless of how reversible it looks. The first is personal data: if the action moves personal data to a new destination, treat it as Class 2 at minimum. The second is asymmetry of harm. A wrong tag costs a correction. A wrong message to a regulator costs a relationship, and averages hide that.
The employment row in Class 3 is not a stylistic choice on our part. Under the EU AI Act, systems used for recruitment, task allocation, performance evaluation, promotion and termination are classified as high-risk in their own right, which is covered in the legal section below. If your AI employee's job description touches any of those, the class is decided for you before you reach the undo test.
This is also what resolves the objection we opened with. The commenter working under NIS2 was effectively told that everything had to be Class 3, while the same organisation ran internal tooling that needed no gate at all. A per-action classification lets both statements be true at once; a single autonomy setting never can.
The security literature supports the same shape. OWASP's LLM06:2025 Excessive Agency defines the risk as damaging actions performed in response to unexpected, ambiguous or manipulated model output, and traces the root cause to excessive functionality, excessive permissions or excessive autonomy. Its named mitigations read like a control list for the table above: minimise extensions and their permissions, execute in the user's context, require user approval for high-impact operations, and enforce authorisation in the downstream system rather than trusting the model to enforce it, which OWASP calls complete mediation.
The onboarding runbook: day 0 to day 30
Here is the procedure we would run, in order. It maps one-to-one onto the five records, and it is deliberately slower at the start than most vendor quickstarts, because the expensive failures all come from granting scope before the function is defined.
Before the first step, settle who does each one. In most companies these tasks sit with three different groups. The business owner writes the job description and takes the baseline, identity and platform engineering issue the credential and the scopes, and the owner's manager or a risk function signs off the class assignments. Thirty days is achievable when those three are named in advance and unachievable when the runbook has to discover them. The most common reason a "one-month pilot" becomes a four-month one is not the technology; it is waiting for a scope grant from a team that was never told it was on the critical path.
- Write the job description (day 0). One function, one trigger, one output, one owner. Include the escalation condition, meaning the circumstances in which the agent must stop and hand to a human, in the same sentence, and name the escalation path it hands to. If you cannot state the escalation condition, you do not understand the function well enough to delegate it.
- Take the baseline (day 0-2). Volume, cycle time, error rate, cost. Store it somewhere that is not a spreadsheet on the owner's laptop.
- Classify every action the function requires (day 2). Use the undo test. You will usually find one Class 3 action hiding inside a job everyone described as routine: approving a supplier, deleting a duplicate record, changing a payment detail. That discovery is the point of the exercise.
- Issue the identity (day 3). Its own credential, its own name, an owner attached, and an expiry date set now rather than "later". Do not reuse an existing automation account, however convenient. An identity shared between two agents makes the fifth record unrecoverable.
- Grant the minimum scope, action-level (day 3-5). Read before write. One system before three. Where the platform supports scoping by action rather than by connection, use it. Read tickets is a different grant from close tickets. Many platforms do not support it, and the fallback is to split the agent's access across more than one identity: one credential that can only read, a second that can write the narrow set of things it is allowed to write, and the write credential held behind whatever approval layer you have. It is uglier than a single scoped grant and it survives an audit, which a single all-purpose connection does not.
- Set the money limit (day 5). A per-agent ceiling in currency, not tokens, with an alert threshold below it. Token budgets are unintelligible to the person who has to approve them, and the cost failures reported in this market are budget failures, not technical ones.
- Run in observe mode (day 5-12). The agent proposes; a human executes; everything is logged. Our own gateway rollout model is Observe first. Enforce second. Optimize third., and the reason observation comes first is that you cannot write sensible rules for traffic you have never seen. A week of proposals tells you the real error distribution, which no vendor benchmark will.
- Turn on Class 1 actions only (day 12). Keep the human gate on everything above. Sample the logs daily.
- Promote Class 2 actions behind an approval gate (day 15-25). The gate is not a formality. Record who approved, when, and what they saw. An approval you cannot reconstruct is not a control.
- Hold the thirty-day review (day 30). Compare against the baseline. Decide one of three outcomes: widen scope, narrow scope, or retire. "Leave it running and see" is not one of the three, and it is how orphaned agents are born.
Class 3 actions do not appear in this runbook on purpose. Our recommendation is that no AI employee executes a Class 3 action in its first quarter, whatever the vendor demo showed. If the business case only works with Class 3 automation from week one, the business case is a request to accept an unmeasured risk.
A completed personnel file, filled in
Below is a personnel file written out in full, for a support-triage function. It is an illustrative template composed for this article, not a customer deployment and not a measurement. Every field is one you can fill in tomorrow, and the point of showing it assembled is that the gaps become obvious in a way a checklist never makes them.
Agent name: support-triage-01 Function: Acknowledge, categorise and first-line answer inbound support email for the self-serve product tier; escalate anything matching the escalation conditions. Trigger: New message in the support inbox. Defined output: Every inbound email has, within 10 minutes: an acknowledgement, a category from the fixed taxonomy, and either a first-line answer or an escalation with a reason. Escalation conditions: Any mention of a security incident, data deletion, refund, legal threat, or an enterprise-tier account. Any message the classifier scores below the confidence floor. Any second contact on the same thread.
| Record | Entry |
|---|---|
| Identity | Own non-human identity, no shared service account; owner recorded on the identity itself; credential expires in 180 days and must be re-approved |
| Job description | The function, trigger, output and escalation conditions above, in one page, version-controlled |
| Scope | Read: support inbox, help-centre articles, account status (read-only). Write: ticket create, ticket tag, reply to customer. No write to billing, no write to CRM, no access to payment data |
| Supervisor | Named support operations lead; deputy named; both can suspend the agent without a ticket to IT |
| Record | Every inbound, every classification with its confidence, every reply sent, every escalation, every refusal — retained 12 months, exportable, immutable |
| Action | Class | Control |
|---|---|---|
| Tag and categorise a ticket | 1 | Logged; sampled weekly |
| Draft an internal note | 1 | Logged |
| Send a first-line reply to a customer | 2 | Human approval for the first 30 days, then approval only below the confidence floor |
| Escalate to a human queue | 1 | Logged |
| Issue a refund | 3 | Not granted. Escalation only |
| Close a ticket without reply | 2 | Human approval, always |
Cost ceiling: a fixed monthly figure in currency, alert at 70 percent, hard stop at 100 percent with escalation to the owner rather than silent failure. Review date: 30 days, then quarterly. Exit conditions: the function is retired, the vendor is replaced, the owner leaves without a successor, or two consecutive quarterly reviews show no measurable improvement over the baseline.
Read that file as a manager rather than an engineer and something becomes clear: the interesting decisions are all in the second table. The identity and scope rows are configuration. The class assignments are policy, and they are the ones a business owner is qualified to argue about.
A template always looks tidier than the deployment it describes, so it is worth saying where this one usually goes wrong. The row that fails first is the escalation condition, because the list above is closed and real inboxes are not: the first message that is angry but not legal, or urgent but not enterprise-tier, will not match any condition and will be answered by the agent. The remedy is not a longer list. It is a confidence floor that catches the unclassifiable, plus a standing instruction that anything the agent could not confidently place goes to a human. The second row that fails is the credential expiry, which nobody re-approves at day 180 because the reminder goes to an individual rather than a rota.
What an AI employee cost actually consists of
An AI employee cost is not one number, and the sticker price is usually the smallest line. Published seat pricing gives you a floor: on Microsoft's own Microsoft 365 Copilot Business page, fetched on 30 July 2026, the add-on is listed at $18.00 per user per month paid yearly (promotional, against a standard $21.00), with Business Premium bundled with Copilot at $32.00 per user per month. That is the assistive tier. Whatever you pay for an AI employee, you pay more than that per person-equivalent, because a function costs more than a seat.
The full picture has six line items, and four of them are invisible at purchase:
| Line item | What drives it | Who usually forgets it |
|---|---|---|
| Platform or seat fee | Users, agents or functions | Nobody |
| Consumption | Runs per month × steps per run × model cost per step | Everybody, until the first invoice |
| Integration | Number of systems it writes to, and whether they have real APIs | Finance |
| Human review time | Class 2 volume × minutes per approval | Everybody |
| Maintenance | Prompt and policy upkeep, connector changes, model deprecations | Engineering, optimistically |
| Exit | Rebuilding the function elsewhere, or bringing it back in-house | Everyone, always |
Consumption is where the surprises live. Fortune reported in May 2026 that Uber's chief technology officer had told The Information the company burned through its entire 2026 AI coding tools budget in four months, and that Microsoft was reportedly cancelling most of its direct Claude Code licences. Those are engineering tools rather than AI employees, and the reporting is second-hand, but the mechanism transfers exactly: per-run costs scale with usage, usage scales with success, and a budget set as an annual line item does not notice until it is gone. This is the argument for setting the ceiling in currency at the agent level on day five rather than reading about it in a quarterly review.
Human review time is the line that decides whether the business case survives contact. If a Class 2 approval takes 90 seconds and the function runs 400 times a month, you have created ten hours of monthly work while removing something else. That is often still a good trade. It is never a good surprise.
Since nobody can give you a credible total, build the arithmetic yourself and fill it with your own numbers: (platform fee) + (runs per month × steps per run × cost per step) + (Class 2 runs × minutes per approval × loaded hourly rate) + (amortised integration and maintenance), divided by units of the function completed, compared against the baseline you took in step two of the runbook. Every term in that expression is knowable after one month of observe mode, and none of it is knowable from a pricing page. The two levers that actually move the review-time term are the confidence floor at which the agent stops for a human and the boundary between Class 1 and Class 2. Both belong on the quarterly review agenda, and neither should move until the correction rate justifies it. We have gone into the wider version of this arithmetic in our analysis of what enterprise AI implementation actually costs beyond the licence.
The review record: how you tell whether it is working
The thirty-day review has one job: compare the agent against the baseline you took before it existed, and decide whether to widen, narrow or retire. Everything else in the meeting is conversation. Formalise it, because the alternative is the most common end state we see: an agent that quietly keeps running because nobody owns the decision to stop it.
Four measures are worth carrying into every review, and they are deliberately boring:
- Deflection. What share of the function no longer reaches a human at all? This is the number the purchase was justified on.
- Escalation quality. Of the items it escalated, how many should it have handled? And, more importantly, how many that it handled should it have escalated? The second figure is the risk number and it is the one people skip.
- Correction rate. How often did a human undo or rewrite the agent's output? A rising correction rate on stable input volume is drift.
- Cost per completed unit. Total spend, including review time, divided by units of the function completed. Compare against the baseline, not against last month.
When the review says "no measurable improvement", the next question is whether you bought a bad agent or defined a bad function, and there is a cheap test. Look at the escalations. If the agent escalated a lot and escalated correctly, the function is fine and the agent is under-capable. Narrow the scope to the part it handled well and re-review. If it escalated rarely but the correction rate is high, it is over-confident, and the fix is the confidence floor rather than the vendor. If the escalations are scattered with no pattern, the function itself was never one job, and no agent will fix that. Only the first of those three is a reason to change product.
Give the review a written outcome and file it against the agent. That file is the appraisal history, and it is what makes the twelve-month conversation possible: an agent with four quarterly reviews on record is a managed asset, and an agent with none is technical debt with a login.
Set expectations accordingly. A Gartner survey of 110 heads of HR published on 27 July 2026 found that while 95 percent of organisations had implemented AI in some capacity over the previous year, only one in five had realised significant or transformational value. McKinsey's survey adds a sharper detail: explainability was the second-most-reported AI risk and was not among the most commonly mitigated. Reviews that find no measurable improvement are the normal case, not the embarrassing one. What matters is whether the process catches it. For the specific reasons pilots stop short of production, see our earlier analysis of why AI pilots stall.
Offboarding: the part almost nobody writes down
Offboarding an AI employee means revoking its credentials, closing its scopes, transferring or archiving its record, and confirming that the function it owned has somewhere to go. It is the single most neglected step in the entire lifecycle, and it is ranked first in the OWASP Non-Human Identity Top 10 for a reason: a credential nobody remembers issuing, attached to an agent nobody owns, is a standing invitation.
The standards bodies already say this out loud. The NIST AI Risk Management Framework has four functions (Govern, Map, Measure and Manage), was published in January 2023, and is currently under revision. Its companion Playbook sets out GOVERN 1.7: "Processes and procedures are in place for decommissioning and phasing out of AI systems safely and in a manner that does not increase risks or decrease the organization's trustworthiness." Almost nobody writing about the digital workforce mentions it, and it is the checkpoint an auditor will find first.
There are four events that should trigger the runbook, and only one of them is the obvious one:
- The function is retired. Expected, planned, easy.
- The vendor is replaced. The agent goes, the function stays, and the risk is that the old credentials stay too because the migration project ends when the new thing works.
- The owner leaves. This is the quiet one. Ownership must transfer to a named successor on the leaver's last day, or the agent goes dormant. An agent whose supervisor left is an unsupervised agent, whatever the config says. We argued the general form of this in our piece on turning personal prompts into owned company assets. Ownership that dies with the person was never ownership.
- A review says retire. The outcome the thirty-day and quarterly reviews are allowed to produce.
The steps themselves are short, and the order matters:
- Suspend before you delete. A suspended agent stops acting and keeps its history. A deleted one may take the history with it. If the platform has no suspend state, and plenty do not, revoke the credentials first, confirm the log export has actually landed somewhere you control, and only then delete the identity. Deleting first and exporting second is how organisations discover their audit trail lived inside the thing they just removed.
- Revoke every credential and every token, including the ones in the connector layer. Long-lived secrets are their own entry in the OWASP list; a rotated key that was never revoked is still a key.
- Close the scopes at the downstream system, not just in the agent's own configuration. Complete mediation cuts both ways. If the ticket system enforces the permission, the ticket system is what must remove it.
- Export and retain the action record for the period you already decided in the prerequisites. Under the EU AI Act, deployers of high-risk systems must keep automatically generated logs "for a period appropriate to the intended purpose of the high-risk AI system, of at least six months" (Article 26(6)). Six months is a floor imposed on a subset of systems, not a ceiling or a target.
- Reassign the function explicitly. Back to a human queue, to another agent, or formally dropped. An offboarding that leaves the work homeless produces a shadow process within a fortnight.
- Record the closure against the agent's file, with the date and the reason. That is the last entry in the personnel file, and it is what makes the next audit answerable.
Run that list once and you will find something unexpected: usually a scope granted for a trial that never expired, or a credential shared with a second agent. That discovery is the value.
Common mistakes we see in the first ninety days
Six mistakes account for most of the disappointment, and all six are organisational rather than technical.
Buying a platform before defining a function. The demo is impressive because the demo has a defined function. Yours does not yet. Writing the job description first costs a morning and changes what you buy.
Handing the agent a human's credentials. It works immediately, which is the problem. The audit trail now attributes machine actions to a person, and every later investigation starts from a false premise.
Treating autonomy as a product setting rather than a per-action decision. A single dial cannot express "tag freely, never refund".
Setting budgets in tokens. Nobody who has to approve the spend understands tokens, so nobody governs them.
Skipping the baseline. Without a before-number the review becomes a debate about impressions, and impressions favour whoever is most invested.
Letting scope creep in through convenience. Each widening is individually reasonable. It needs CRM read to answer properly. It needs write to close the loop. Six reasonable widenings later, the blast radius is unrecognisable. Re-derive scope from the job description at each quarterly review instead of accumulating grants.
A seventh is less a mistake than a category error: expecting the AI employee to cut headcount in its first quarter. Gartner's July 2026 survey found 22 percent of CHROs reported at least one business leader who had stopped hiring for entry-level roles because of AI automation. That is a real effect, but it shows up as slower hiring rather than reduced staff, and only once the function has run long enough to trust.
What the law already requires of you
If your AI employee touches recruitment, task allocation, performance evaluation, promotion or termination, EU law already classifies it as high-risk, and the obligations are specific rather than aspirational. Annex III point 4 of the EU AI Act covers "Employment, workers' management and access to self-employment". In the regulation's own terms that reads on systems used to filter job applications and evaluate candidates, and on systems used to decide terms of work-related relationships, allocate tasks, or monitor and evaluate a worker's performance and behaviour.
Four deployer obligations map directly onto the personnel file:
| Provision | What it says | Which record it maps to |
|---|---|---|
| Article 26(2) | Deployers "shall assign human oversight to natural persons who have the necessary competence, training and authority, as well as the necessary support" | Supervisor |
| Article 26(6) | Deployers keep the automatically generated logs "for a period appropriate to the intended purpose... of at least six months" | Record |
| Article 26(7) | Before putting a high-risk system into service at the workplace, employers "shall inform workers' representatives and the affected workers that they will be subject to the use of the high-risk AI system" | Job description and scope |
| Article 14(4)(e) | Human overseers must be able "to intervene in the operation... or interrupt the system through a 'stop' button or a similar procedure that allows the system to come to a halt in a safe state" | Supervisor |
Article 26(7) deserves more attention than it gets in vendor material. If you deploy a high-risk AI system at the workplace, telling your staff is not a communications choice. It is a legal obligation, and it lands before go-live rather than after. Any plan that treats the rollout as a quiet efficiency measure is already non-compliant on that point.
Article 14(5) is worth reading for a different reason: for one category of high-risk system it requires that no action be taken unless the output has been "separately verified and confirmed by at least two natural persons with the necessary competence, training and authority." That is dual control, written into statute. Our Class 3 is not a novel invention; it is the pattern the regulator already reached for when the stakes were high enough.
Two voluntary frameworks fill in the rest of the management scaffolding. The NIST Playbook gives you GOVERN 3.2, policies that "define and differentiate roles and responsibilities for human-AI configurations and oversight of AI systems", alongside the decommissioning requirement quoted earlier. ISO/IEC 42001:2023, described by ISO as the world's first AI management system standard, specifies requirements for establishing and continually improving an AI management system in organisations that provide or use AI-based products and services. Neither is a certification you need before deploying a support-triage agent. Both are useful as a shopping list when a customer's security questionnaire arrives. For the timeline detail on what applies when, see our guide to EU AI Act compliance for AI deployers.
If you want to hear this discussed by people running it at scale rather than selling it, this recorded conversation with the CIO and CISO of a 5,500-person company covers identity, scope and observability for a fleet of production agents:

Where this is still uncertain
Several parts of this argument are less settled than the confident tone of the category suggests, and it is worth naming them.
The category boundary is contested, and our test is a proposal, not a standard. The Personnel File test is a framework we find useful for making purchase and deployment decisions legible. It is not an industry definition, no standards body has endorsed it, and a reasonable person could argue that the fifth record, a reconstructable action log, belongs inside the first four rather than beside them.
The delegated-authority model may simply be right. The agent-as-app objection is strong, and in a world where every downstream system supports fine-grained, just-in-time authorisation, most of the identity argument dissolves. We do not think that world exists yet in the average enterprise, but the direction of travel is towards it, and an article written in 2028 might weight this differently.
We have not benchmarked any AI employee product. There is no first-hand test behind this piece, no measured deflection rate of our own, and no vendor comparison. The numbers here are other people's, cited and dated.
The employment-effect research is early and moves. The Stanford paper cited in the FAQ revised its headline figure between its 2025 versions, where the widely quoted 13 percent became 16 percent. That is what should happen with new data, and exactly why single statistics about AI and jobs should be held loosely.
Cost data in this market is bad. Vendors publish seat prices and not consumption curves, third-party AI employee cost comparisons are mostly lead generation, and the most useful cost evidence in this article is a second-hand report of one company's budget overrun. Treat any total-cost claim, favourable or not, as unverified until you have run the function for a quarter.
The regulatory timeline is not fixed. The EU AI Act phases in over several years and has already been through a simplification package; dates and scope have moved before.
How LeapForce approaches the personnel file
LeapForce builds the layer this article describes: one controlled place where every AI tool, connector, model and agent runs under an identity, a policy, a budget and an audit trail. In our own product, a coworker is created through a lifecycle of Build → Scope → Review → Share → Improve. The Scope step issues the agent its own non-human identity with a named owner, the minimum connector scopes it needs, a model policy and a budget. The Review step records who approved it before it goes live, so the question "who allowed this agent?" always has an answer. Our rollout model for the gateway itself is Observe first. Enforce second. Optimize third. That is why the runbook above spends a week watching before it writes a single rule. We label per-capability build status publicly, and the design intent that budgets are expressed in dollars rather than tokens should be read against those labels rather than assumed shipped.
One clarification and then the CTA. LeapForce does not sell an AI employee, and we do not build your support-triage or finance agents; we build the governed layer they run inside. If you are choosing a platform to build the agent itself, the personnel file is a specification you can hand to any vendor, including one that is not us.
Frequently asked questions
No. The dividing line is write access. A chatbot answers inside the conversation and the state of your business is unchanged when the window closes; an AI employee changes something outside the conversation: it creates the ticket, sends the email, updates the record. That is why the two need different controls: a wrong chatbot answer costs a correction, while a wrong write costs an undo, and some writes cannot be undone at all. If the software cannot write to a system of record, it is an assistant with good marketing.
An event or a schedule triggers it. It reads context from the systems it has been granted access to. A model plans the next action from the goal and the context. Before each action runs, an authorisation check asks whether this identity is permitted this action on this resource; actions above the permitted class stop for a human approval. Permitted actions execute against the target system through a scoped connector. Every step, including the refusals, is written to a log with the agent's identity and the authorisation that allowed it. The last part is what separates a governed AI employee from a script with a language model in the middle.
There is no single figure, and any article that gives you one is guessing. What you can price is the structure: a platform or seat fee, plus consumption that scales with runs and steps per run, plus integration work, plus human review time on every action that needs approval, plus ongoing maintenance, plus the eventual cost of exit. For an anchor at the assistive end, Microsoft lists Microsoft 365 Copilot Business at $18.00 per user per month paid yearly as of 30 July 2026. A function-owning agent costs more than a seat, and the variable component is the one that surprises people. Fortune reported in May 2026 that one large company exhausted an annual AI tooling budget in four months.
Sometimes, at the margin, and later than the pitch suggests. Gartner's July 2026 survey of 110 heads of HR found 22 percent reported at least one leader who had stopped hiring for entry-level roles because of AI automation. That is a real effect, but it shows up as slower hiring rather than as redundancies. Research from the Stanford Digital Economy Lab (Brynjolfsson, Chandar and Chen, November 2025) found a 16 percent relative decline in employment for workers aged 22 to 25 in the most AI-exposed occupations, concentrated where AI automates rather than augments. Our practical advice is to justify the purchase on the function, not on a headcount line, because the headcount effect arrives after the function is trusted and the function takes a quarter to trust.
The named human owner, and if there is no named owner then whoever authorised the scope. Software cannot hold accountability. It cannot be disciplined, sued or given a duty of care, so accountability sits with a person by construction. The EU AI Act makes this explicit for high-risk systems in Article 26(2), which requires deployers to assign human oversight to natural persons with the necessary competence, training and authority. Operationally, the answer to "who is accountable" is only real if you can reconstruct what happened, which is why the action record is one of the five personnel-file entries rather than a nice-to-have.
Its own, in almost every case where it runs on a schedule or on an event with nobody in session. Sharing a person's credentials makes the audit trail lie, because machine actions appear as that person's actions, and it grants the agent everything the human can do, including the destructive parts. The honest exception is an agent that acts strictly on behalf of a specific user during that user's session, where the delegated-authority model is cleaner: the agent holds its own application identity but the permission check is against the user. The mistake is mixing the two, most often by pasting a long-lived human token into a scheduled job.
Plan for about a month to a supervised production state for one narrow function, and a quarter before you trust it unsupervised for anything reversible. The runbook in this article is thirty days: define the function and take a baseline in the first two, classify the actions and issue identity and scope in the first week, run in observe mode for a week, enable low-risk actions, then promote higher-risk actions behind an approval gate, and review at day thirty. Vendor quickstarts are much faster because they skip the classification and the baseline, which are the two steps that make the day-thirty review meaningful.
In the EU, if the system is high-risk and deployed at the workplace, yes. Article 26(7) of the AI Act requires employers to inform workers' representatives and the affected workers before putting it into service. Systems used for recruitment, task allocation, performance evaluation, promotion or termination fall under Annex III point 4 and are high-risk by classification. Outside those categories and outside the EU the legal duty varies, but the practical argument is the same everywhere: a digital workforce discovered by staff rather than announced to them starts every subsequent conversation from a deficit.
Technically yes, and it is usually a bad first move. Each additional function widens the scope, adds actions in higher supervision classes, and makes the review ambiguous. When deflection improves and the correction rate worsens, you cannot tell which function moved. Our rule is one function per agent until each has a clean quarterly review, then consolidate deliberately if the scopes genuinely overlap. Bundling early is the most common route to an agent whose blast radius nobody can describe in a sentence.
Suspend it rather than deleting it, so the history survives. Revoke every credential and token, including any held in the connector layer. Close the scopes at the downstream systems, not only in the agent's own configuration. Export and retain the action record for your decided retention period. For high-risk systems under the EU AI Act that floor is at least six months. Reassign the function explicitly to a human queue, another agent, or formally drop it. Then record the closure against the agent's file with the date and reason. OWASP ranks improper offboarding as the top non-human identity risk for 2025, ahead of secret leakage, which tells you how often those steps are skipped.
Five questions, one per record. Can this agent hold its own identity with a named owner and an expiry, or does it need a shared account? Can scopes be granted per action rather than per connection? Can I require human approval on a specific action rather than switching the whole agent to supervised mode? Can I set a spending ceiling in currency per agent, with an alert below the stop? And can you show me, for an action taken last month, who authorised it and on whose behalf, exported rather than shown on a dashboard? A vendor who answers the fifth question well has usually got the first four right.
Yes, with a narrower definition of the word workforce. A small company can run one or two function-owning agents well, and the personnel file scales down without losing much. The owner is the person who runs the function, the review is a calendar entry, the record is whatever your platform exports. What does not scale down is the classification step: a five-person company still needs to know which actions cannot be undone, and it has fewer people available to catch the ones that go wrong. The realistic small-company pattern is one Class 1 agent that works, not five agents that need supervising.
Ready to Govern Your AI?
Talk to LeapForce — one controlled layer for every AI tool, connector, model, and agent.
Comments