Vertical AI Agents: Domain Knowledge Is the Cheap Half

A vertical AI agent is an AI system built to do the work of one industry — recruiting, claims, underwriting, clinical documentation — inside that industry's own

A vertical AI agent is an AI system built to do the work of one industry — recruiting, claims, underwriting, clinical documentation — inside that industry's own systems of record, rather than answering questions about it. Domain knowledge is the cheap half of that.

Our position at LeapForce is that the vocabulary, the process order and the trained-on-your-corpus part have all commoditised faster than anyone selling them expected. What has not commoditised is the second half: a vertical AI agent is only "vertical" to the degree it has been granted authority inside the record — the CRM opportunity, the applicant tracking file, the chart, the ledger line. That grant is the product. It is also the risk, the compliance surface and the reason pilots stall, and almost nobody's buying guide asks about it.

Someone on Hacker News put the practical version of this better than any vendor page. In September 2025, a user posting as Hoshang07 asked for a "non-career ending way" to let agents touch internal structured data, opening with a line most platform engineers will recognise: "i wouldn't trust an ai agent anywhere near my warehouse with raw sql". What he wanted was not a smarter model. He wanted a defined view, a fixed set of callable operations, and a boundary the agent could not talk its way past.

The short answer: Choose a vertical AI agent on what it is allowed to write to your system of record, under whose identity, with whose approval, and what the resulting log proves — not on how well it speaks your industry's language, because every serious contender now does.

Last updated: July 30, 2026.

Six rungs from advising a user to acting outside the company, each with its own control and reversal cost

The Record Authority Ladder: what an agent may do to the system of record, and what each rung costs to undo.

One disclosure before the argument. We have not run a controlled trial of vertical agents inside a customer's CRM or EHR and we are not going to pretend otherwise, so every number below is somebody else's, named and linked. What we bring is the deployment-layer view: we build the access, policy, cost and audit layer that sits under agents like these, and that vantage point is what the argument is made from.

What a Vertical AI Agent Actually Is

A vertical AI agent is a system that carries out a bounded piece of one industry's work end to end: it reads from that industry's systems, applies that industry's decision rules, and produces the outcome rather than a description of the outcome. A recruiter wants a shortlist in the applicant tracking system, not a screening dashboard.

The category travels under other names too: industry-specific AI agents, domain-specific agents, verticalised intelligence. The labels matter less than the property they all point at. Three things distinguish one from a general assistant pointed at an industry problem. It knows the process order — that a prior authorisation goes out before the procedure is scheduled, not after. It speaks in the objects the industry's software actually stores, so "requisition", "encounter" and "opportunity" are records with fields, not words in a prompt. And it holds credentials into those systems, which is where the interesting part starts.

That third property is the one this article is about, because the first two have collapsed in price. Frontier models call tools reliably now. Retrieval over a domain corpus is a weekend. Whatever moat "verticalised intelligence" described three years ago, it is not sitting in the model any more.

The market data reads the same way. Gartner predicted in August 2025 that 40% of enterprise applications would be integrated with task-specific AI agents by the end of 2026, up from less than 5% at the time of writing. When four in ten business applications ship with an embedded task agent, "has an agent that understands your industry" stops being a differentiator and becomes a checkbox. What remains scarce is trustworthy authority inside the record.

Bessemer Venture Partners made the investment case for the category in September 2024, noting that the top 20 US public vertical SaaS companies had a combined market capitalisation of roughly $300 billion and predicting vertical AI would end up at least ten times that. Their supporting observation is the one worth keeping: LLM-native companies in these categories were pricing at around 80% of the contract value of the incumbent core system while growing about 400% year over year. Buyers are not paying a premium for a smarter chatbot. They are paying near-incumbent prices for something that does the work — which means it has to be let into the place the work lives.

Vertical versus horizontal, stated precisely

The usual framing, that horizontal agents are broad and vertical agents are deep, is true but not useful at purchase time. The sharper distinction is about scope of authority.

Horizontal agentVertical AI agent
Where it operatesGeneric surfaces: mail, docs, chat, searchThe industry's system of record
What it must knowHow to use toolsThe record model and the process order
Worst realistic failureA wrong answer a person readsA wrong record a downstream process acts on
Who notices a mistakeThe user, immediatelyAn auditor, months later
Governing questionCan it see confidential data?Can it change a record, and who authorised that?

That last row is the whole article. A horizontal agent's failure mode is epistemic; a vertical agent's failure mode is transactional. Once an agent's output becomes a row that a billing engine, a payroll run or a clinical workflow consumes, the interesting question stopped being accuracy and became authority.

The Record Authority Ladder

Every vertical agent sits on exactly one of six rungs, defined by what it may do to the system of record and how expensive that is to undo. We call this the Record Authority Ladder, and it is the single diagnostic we would run before comparing any two vertical AI agent platforms. It takes about an hour with the process owner and one person who knows the system's permission model.

It is a narrower instrument than it looks. We have argued more generally that the line worth drawing across all agents is the write-access line — the point at which a system stops describing and starts changing things. The ladder below is that argument applied to one specific object: not a filesystem, not a sandbox, but the record your industry runs on and your regulator will ask about.

RungWhat the agent doesReversal costThe control that actually binds
L0 AdviseReasons only over what the user gives itNoneData handling on the way in
L1 ReadQueries the record; changes nothingNone, but disclosure is permanentRow and field scoping; where answers may go
L2 DraftProduces an artifact held outside the recordDelete the draftProvenance labelling on the draft
L3 Write with approvalStages a record change a named human commitsReverse the commitApproval identity and the diff shown
L4 Write unattendedCommits record changes inside a scopeReverse plus reconciliationScope, rate limits, blast-radius caps
L5 Act outwardTriggers an external effect: offer sent, claim filed, payment releasedOften impossiblePre-commit policy check and a durable log

Three rules make the ladder useful rather than decorative.

The rung is a property of the deployment, not the product. The same vendor's recruiting agent is L2 at one company and L4 at another, because one of them wired it to write dispositions. Vendor comparison tables cannot tell you your rung. Only your integration can.

Each rung has its own control set, and controls do not carry upward. Field-level read scoping is the right control at L1 and is nearly irrelevant at L5, where the binding control is a policy check that runs before the outbound action commits. Teams that build careful L1 controls and then jump to L4 have not built 80% of what they need; they have built something else.

Most programmes are bought at L4 and governed at L1. This is the mismatch we would look for first, and it follows from how the two documents get written. The business case is written on unattended execution — that is where the labour saving is — while the security review examines what data the agent can see, because read access is what security reviews have always examined. The gap between the two is where the cancelled projects live.

That last claim is not ours alone. Gartner predicted in June 2025 that over 40% of agentic AI projects would be cancelled by the end of 2027, naming three causes: escalating costs, unclear business value, and inadequate risk controls. The same release put a number on the vendor noise: Gartner estimated only about 130 of the thousands of agentic AI vendors are real, and separately named the practice of rebranding assistants, RPA and chatbots without substantial agentic capability "agent washing".

Running the ladder assessment

Pick one workflow you are considering handing to a vertical agent and answer five questions in order. Stop at the first "yes" you cannot control.

  1. Does it read the record? If yes, list the tables and fields — actually list them. "Access to the CRM" is not an answer.
  2. Does anything it produces land in the record? If yes, does it land as a draft object or a live field?
  3. Does a named human commit the change? Named, not "someone on the team". A shared service account is a "no".
  4. Can it commit changes without a human in the loop? If yes, what is the maximum number of records one run can touch before something stops it?
  5. Can any of its actions leave the company? Email to a candidate, submission to a payer, a payment instruction, an API call to a partner. If yes, you are at L5 and you need a pre-commit policy check, not a post-hoc review.

Write the answer down as a rung. Then check what your shortlisted platform actually enforces at that rung, rather than what its marketing describes at L1.

Why the Rung Matters More Than the Vertical

The strongest available evidence that domain fit is not the differentiator comes from the vertical where domain fit was supposed to matter most. In a randomised trial published in NEJM AI on 26 November 2025, Lukac and colleagues assigned 238 outpatient physicians across 14 specialties to one of two ambient AI scribe products or usual care. Both products were purpose-built for clinical documentation. Both were trained on medical language. The results diverged anyway: Nabla users showed a 9.5% decrease in time-in-note against control (95% CI −17.2% to −1.8%, P=0.02), while the other product showed no significant change (−1.7%, P=0.66).

Read that carefully, because it cuts both ways. Two agents in the same vertical, on the same task, with the same clinical vocabulary, produced materially different operational results — so "it's built for healthcare" predicted almost nothing. The trial also recorded that clinicians rated clinically significant inaccuracies as occurring "occasionally" for both products, on a five-point scale (2.7 and 2.8). An agent that is right most of the time and occasionally materially wrong is precisely the profile that makes the rung question load-bearing: at L2 an occasional error is a draft someone fixes, and at L4 it is a note in a chart.

Be careful how far that evidence stretches, though. Ambient scribes are a low-rung deployment. The physician signs the note, so they sit at L3 by design, and the trial measured documentation time rather than autonomous decision quality. What it establishes is narrow but real: within one vertical, on one task, purpose-built domain products diverged sharply on the outcome that mattered. It does not establish anything about unattended agents, because nobody has run that trial.

The scale picture supports the same reading. McKinsey's State of AI survey, published 5 November 2025 and fielded to 1,993 respondents in 105 nations, found 23% of organisations scaling an agentic AI system somewhere in the enterprise and a further 39% experimenting — but in any given business function, no more than 10% reported scaling agents. Enthusiasm is broad; production depth is thin. Notably for this argument, McKinsey found agent use most widely reported in technology, media and telecommunications, and healthcare: three sectors where industry-specific AI agents run against systems of record that are unusually consequential and, in healthcare's case, unusually regulated.

The same survey supplies the discipline. Fifty-one percent of respondents at organisations using AI said they had seen at least one negative consequence, with nearly a third reporting consequences from AI inaccuracy. More telling, explainability was the second-most-commonly-reported risk yet was not among the most commonly mitigated. That is a governance gap with a name: organisations are experiencing a problem of not being able to reconstruct why the system did something, and are not investing in the ability to reconstruct it.

The sales vertical is running the experiment in public

Gartner's most recent statement on the topic, published 28 July 2026, predicts AI agents will outnumber sellers ten to one by 2028 — while fewer than 40% of sellers will say agents improved their productivity. VP Analyst Dan Gottlieb's framing is the sharpest line anyone has written about vertical deployment: "If those systems are fragmented, the agents will scale the fragmentation."

That is the CRM version of the ladder argument. A sales agent granted L4 authority over opportunity records in a CRM where ownership, stage definitions and field hygiene are already contested does not fix the contest. It multiplies it, at machine speed, with worse attribution. The prerequisite is not a better sales agent. It is knowing which records the agent may write and being able to prove afterwards which writes were its.

This recorded panel is the most useful practitioner account we found of an organisation actually operating at that scale — Cvent's CIO and CISO describing identity, scope and observability across a large agent estate.

Play video

Your Vertical Already Wrote the Rules Down

Here is the part that vertical AI agent roundups consistently skip: in most regulated verticals, the constraint on what an agent may do to the record is not a design preference. It already exists in law or rule, it predates agents, and it is enforceable today.

VerticalThe recordThe rule that already bindsWhat it constrains
RecruitingApplicant tracking systemNYC Local Law 144 of 2021; EU AI Act Annex III(4)Screening and selection; audit and notice duties
HealthcareElectronic health recordHIPAA minimum necessary standardWhich records the agent may see at all
Lending and insuranceOrigination and policy systemsEU AI Act Annex III(5)(b)Creditworthiness evaluation and scoring
Public benefitsEligibility systemsEU AI Act Annex III(5)(a)Eligibility determinations, including healthcare services
SalesCRMContractual and data-protection commitmentsWhere customer data goes and what may leave
Four verticals mapped to their system of record and the rule that already governs an agent writing to it

The constraint on a vertical AI agent is usually older than the agent.

Recruiting. New York City's Department of Consumer and Worker Protection states plainly that Local Law 144 of 2021 "prohibits employers and employment agencies from using an automated employment decision tool unless the tool has been subject to a bias audit within one year of the use of the tool, information about the bias audit is publicly available, and certain notices have been provided to employees or job candidates." Enforcement began 5 July 2023. An agent that ranks or filters candidates is squarely inside that definition, and no amount of "it's just a copilot" framing moves it out. Note the scope precisely: this is a city ordinance, and its reach is defined by the jobs and candidates it covers rather than by where your engineering team sits — which is exactly the kind of detail a vertical agent rollout gets wrong when it treats one applicant tracking system as one jurisdiction.

Europe adds a second layer. Annex III of the EU AI Act classifies as high-risk those AI systems "intended to be used for the recruitment or selection of natural persons, in particular to place targeted job advertisements, to analyse and filter job applications, and to evaluate candidates", and separately those used "to make decisions affecting terms of work-related relationships, the promotion or termination". The same annex covers systems used "to evaluate the creditworthiness of natural persons or establish their credit score" and those evaluating "eligibility of natural persons for essential public assistance benefits and services, including healthcare services". If your vertical agent does one of those things, high-risk obligations attach to the deployment, not to the vendor's marketing category. We have written a fuller treatment of what that means for deployers in our earlier analysis of EU AI Act compliance.

Healthcare. The HIPAA minimum necessary standard, at 45 CFR 164.502(b) and 164.514(d), requires covered entities to "take reasonable steps to limit the use or disclosure of, and requests for, protected health information to the minimum necessary to accomplish the intended purpose", and requires policies that "identify the persons or classes of persons within the covered entity who need access to the information". Read that as an engineering requirement and it says: a clinical agent needs a defined class with a defined need, and a scope narrower than the human clinician's. There is a carve-out — HHS lists disclosures to or requests by a health care provider for treatment purposes among the situations the standard does not apply to — and it is the single most misread line in this area. It does not mean a clinical agent is exempt because it works in a treatment context. It means a specific class of disclosure is out of scope, and whether your agent's reads fall inside that class is a determination someone has to make and write down, not an assumption to inherit from the vendor's compliance page.

There is also a scoping question the rule does not answer for you. "List the tables and fields" is a reasonable instruction for a CRM and an absurd one for an electronic health record, where the schema runs to thousands of elements. The workable version is to invert it: define the narrow set of retrievals the workflow actually needs — this encounter, these results, this problem list — expose those as named operations, and deny by default. That is the same pattern the Hacker News thread above arrived at independently, and it is the only version that stays true as the schema drifts.

The practical consequence is worth stating in one line, because it reverses the usual buying order: in a regulated vertical, the compliance obligation determines the highest rung you may occupy, and the highest rung you may occupy determines which platforms are even eligible. Domain fit is the last filter, not the first.

What a Vertical AI Agent Is Not

Definition by contrast does more work here than definition by attributes, because three adjacent things are routinely sold under the same label.

It is not vertical SaaS. The vertical SaaS versus vertical AI agents distinction is the one buyers get wrong most often, and it is not subtle once stated. Vertical SaaS is an instrument a domain expert operates: the software presents forms, dashboards and workflows, and a person supplies the judgement and the keystrokes. A vertical AI agent supplies the judgement and the keystrokes, and the human supplies the objective and — if you have designed it well — the approval. The commercial consequence follows directly: you stop buying seats for people to work in a system and start buying outcomes produced inside it. The governance consequence follows just as directly. Seat-based access control assumes a person behind every session, and there is no longer a person behind every session.

It is not a copilot or a plugin. A copilot fires a single scoped action inside an application you are already looking at, and you see the result before it matters. Gartner's staged model puts embedded assistants at stage one and task-specific agents at stage two, and explicitly names the confusion between them "agentwashing". The test is not autonomy in the abstract. It is whether a human is looking at the screen when the change lands.

It is not a fine-tune. A model fine-tuned on your industry's corpus knows more words. It does not thereby acquire an identity in your systems, a scope, an approval path or a log. Every hard problem in this article survives a perfect fine-tune untouched, which is a good reason to treat "trained on domain-specific datasets" as a feature claim rather than an architecture.

It is not robotic process automation with better language. RPA is deterministic: it does the same thing every run, and when the underlying screen changes it breaks loudly. A vertical agent is probabilistic: it adapts when the screen changes, which is the selling point, and it can also decide differently on two identical inputs, which is the part that needs a control. Do not inherit RPA's governance model, which assumed the automation could not surprise you.

Whose Credential Is the Agent Using?

Ask a vertical AI agent vendor which identity their agent authenticates as inside your system of record. The answers cluster into three patterns, and only one of them survives an audit.

Pattern one: a shared service account. Fast to set up, and the reason it is popular. It also means every action the agent takes is attributed to one identity used by many workflows and, usually, by some humans. When someone asks six months later why a candidate was rejected or a claim was coded a particular way, the log says the service account did it. That is not an answer to a regulator and it is not an answer to a customer.

Pattern two: impersonating the requesting user. Cleaner for attribution, since the write shows the human's name. Worse for scope, because the agent inherits everything that human can do, including the parts of the record nobody intended it to reach. It also produces a log that is actively misleading: it records a person performing an action they did not perform, which is a bad artifact to hand an investigator.

Pattern three: a distinct non-human identity per agent, with a named human owner, an explicit scope, an expiry, and its own trail. It is more work up front. It is the only pattern where "who did this, under what authority, and is that authority still valid" has an answer.

Standards bodies have converged on the third pattern. In February 2026, the NIST National Cybersecurity Center of Excellence published a concept paper, "Accelerating the Adoption of Software and Artificial Intelligence Agent Identity and Authorization", which frames the problem as understanding "the potential risks from giving AI agents access to diverse data sets, tools, and applications, and applying appropriate identification and authorization controls to mitigate these risks". The Cloud Security Alliance reached the same place from the threat side in its May 2026 whitepaper on the non-human identity governance vacuum, reporting that only 15% of organisations have high confidence in preventing attacks based on non-human identities and that 47% of such identities had not changed in over a year.

The credential hygiene data is worse than the policy data. GitGuardian's State of Secrets Sprawl 2026, published 17 March 2026, counted 28.65 million new hardcoded secrets added to public GitHub commits during 2025, a 34% year-over-year increase, with AI-service secrets specifically reaching 1,275,105 — up 81%. The report also found roughly 70% of credentials confirmed valid in 2022 were still valid in January 2025, and above 64% a year later. Long-lived, widely copied credentials are the substrate most vertical agents are currently being deployed on top of.

This is the ordinary discipline of machine identity applied to a new consumer, and we have set out the mechanics (owner, scope, expiry, one-step revocation) in our earlier analysis of non-human identity for AI agents. The vertical twist is only that the blast radius is a system of record rather than a sandbox.

The second Hacker News thread worth reading here asked the enforcement question directly. In January 2026, a user posting as amjadfatmi1 asked practitioners what their enforcement point is "that the agent cannot bypass" — and, pointedly, whether they log policy decisions separately from execution, "so you can answer 'why was this allowed?' later". That second half is the one most teams have not built. Recording what ran is standard. Recording what was refused, and why, is what turns a log into evidence.

Four Ways to Get a Vertical AI Agent

There are four routes to a working vertical agent, and they differ far more in what you control than in what they do. The at-a-glance table first, then a uniform block for each.

RouteWhat you getWhat you actually controlTypical failureRealistic starting rung
Buy a vertical AI vendorA finished agent for one workflowConfiguration and the integration scopeIts identity model is fixed and it is not yoursL2–L3
Turn on your incumbent's agentAn agent inside the system you already runVery little beyond enable and role mappingPermissions inherit the seat modelL1–L3
Compose on a horizontal platformConnectors, orchestration, approval gatesScope, identity, policy, logging, costYou own the domain logic, and the maintenanceL1–L4
Build in-houseExactly what you specifiedEverythingTwelve months later you maintain a platformL1–L4

Route 1: Buy a vertical AI vendor

Best for: one high-volume workflow with a clear owner, where the vendor has real depth and you have no appetite to build. What you control: which records the integration touches, which fields it may write, and whether a human commits. Little else. Pros: fastest credible path to value; the domain edge cases are already handled; the vendor absorbs model churn. Cons: the agent's identity model is the vendor's, not yours; your audit trail lives partly in their system; scoping is only as granular as their integration exposes. Cost shape: per-seat, per-workflow or per-outcome subscription, plus integration effort that is usually underestimated because it is a permissions project wearing an integration costume. When this is the wrong pick: when the workflow spans two systems of record, or when your regulator will ask you, not the vendor, to produce the decision log.

Route 2: Turn on your incumbent's embedded agent

Best for: getting an honest baseline before you buy anything, and for L1–L2 work where the answer stays inside the application. What you control: enablement, role mapping, and whatever feature flags exist. Assume the rest is the vendor's. Pros: no new data path, no new vendor, no new credential; the data never leaves a system that already holds it. Cons: permissions almost always inherit the human seat model, which means the agent sees what the role sees; cross-system work is out of reach by construction. Cost shape: an edition upgrade or a per-user add-on, sometimes consumption credits on top. When this is the wrong pick: anywhere you need attribution separate from the human user, or a scope narrower than an existing role.

Route 3: Compose on a horizontal agent platform

Best for: organisations that will run more than one vertical agent and want one control surface across all of them. What you control: identity, scope, connector permissions, approval gates, routing, budgets and the log. That is the whole point of the route. Pros: the governance layer is built once and reused; adding the second and third agent is dramatically cheaper than the first; you can hold L4 agents to the same policy as L1 ones. Cons: the domain logic is yours to write and keep correct, and a platform with excellent controls and no domain depth will lose to a focused vendor on a single workflow. Cost shape: platform fee plus model consumption, with the real cost in the first workflow's design and the ongoing ownership of process changes. When this is the wrong pick: a single workflow, no second use case in view, and no internal owner. Then you are buying a platform to run one thing.

Route 4: Build in-house

Best for: the workflow that is your actual competitive advantage, where the process is the product. What you control: everything, which is both the argument for and the argument against. Pros: exact fit; no vendor between you and your record; the institutional knowledge stays. Cons: you are now maintaining evaluation harnesses, connector permissions, an approval UI, a policy engine and an audit store — none of which is the thing you set out to build. Cost shape: engineering headcount, indefinitely. Treat it as a product line, not a project. When this is the wrong pick: any workflow where a credible vendor exists and the process is not differentiating. Most workflows.

Selection criteria that survive contact. Whichever route you take, four questions separate the shortlist. Can the agent hold an identity distinct from both a shared account and a human user? Can permissions be scoped per action rather than per connection? Does a policy decision get logged separately from the execution? And can you revoke everything the agent can reach in one step when the workflow retires or the owner leaves?

Ask those in writing, and treat a non-answer as an answer. A vendor that cannot describe its identity model in two sentences does not have one you would like. That is not cynicism; identity architecture is the kind of thing engineering teams are proud of when it exists.

One honest complication, because a practitioner will raise it within five minutes: many systems of record do not expose action-level permissions at all. Plenty of enterprise systems offer object and field permissions bound to a role, and nothing finer. Where that is true you cannot scope per action inside the system, and the scoping has to happen one layer out — at the connector or gateway that brokers the call, which is where a narrow set of named operations can be defined and everything else refused. That is more work than a checkbox, and it is the difference between an agent that is limited and an agent that is merely expected to behave.

Worked Example: Walking a Recruiting Agent Up the Ladder

Take one concrete case and move it up the ladder rung by rung, watching what changes at each step. A recruiting agent for a mid-sized employer hiring into New York and the EU: it reads requisitions and applications in the applicant tracking system, and the ambition is to have it screen and progress candidates.

L0 · Advise. A recruiter pastes a job description into a chat and asks for better screening criteria. Nothing touches the applicant tracking system. The only control needed is data handling on the way in: is confidential compensation data going into a third-party model, and where does that model run?

L1 · Read. The agent is connected to the applicant tracking system with read access to open requisitions and applications. Now the scoping question is concrete: which requisitions, which fields, and can it see the demographic fields collected for reporting? The right answer is almost always no to that last one, It is easy to get wrong, because the field sits in the same table.

L2 · Draft. The agent produces a ranked shortlist as an artifact, a document or a proposed list, that lives outside the record until a recruiter acts on it. Add a provenance label, so anyone reading the shortlist next quarter knows it was machine-generated and by which version. A large share of the realistic value sits right here. Pause before climbing further.

L3 · Write with approval. The agent stages a disposition on each candidate; a named recruiter reviews the diff and commits. Two controls become mandatory. The approval must be attributed to a person, not to a role or a shared login. And the reviewer needs to see what the agent proposes to change, not a summary of it.

L4 · Write unattended. The agent sets dispositions itself inside a scope. Now the ladder collides with the law. Screening and filtering applications is exactly what NYC Local Law 144 covers, so a bias audit within the prior year and published results are prerequisites, not follow-ups. Under EU AI Act Annex III(4), the same behaviour is high-risk in Europe. You also need a blast-radius cap — a maximum number of candidates one run may disposition before it halts — because the difference between a bug at L3 and a bug at L4 is the difference between one wrong draft and four hundred wrong records.

L5 · Act outward. The agent sends the rejection email. The action is now irreversible in the way that matters: the candidate has read it. The binding control moves before the action rather than after it: a policy check that runs pre-commit. The log has to record the refusals too, so you can answer why the other three hundred were not sent.

The instructive part of this example is where the value curve and the cost curve diverge. Most of the recruiter's time is saved between L1 and L3. Most of the compliance and engineering cost appears between L3 and L5. A programme that jumps to L4 because the business case was written there frequently discovers it bought the expensive half of the ladder for a marginal share of the benefit.

So what do we actually recommend? Default to L3 and stay there until two conditions hold. The approval step has stopped changing anything, meaning reviewers are accepting essentially every diff, measured rather than felt. And you can produce a log that answers "why was this allowed?" for a sample of past decisions. Only then does L4 buy you something real, because at that point the human in the loop is genuinely ceremonial and the evidence trail exists to catch the cases where it should not have been.

If your business case only works at L4, that is legitimate and common — high-volume claims triage does not pencil with a human on every record. Say so early rather than discovering it at the security review, and budget the L4 controls as part of the case rather than as an overrun: scoped identity, a blast-radius cap, a policy log separate from the execution log, and whatever audit your vertical already demands.

Common failure modes we would look for at each rung. At L1, over-broad reads that quietly include protected fields. At L2, drafts with no provenance that get pasted into the record by a human and become indistinguishable from human work. At L3, approval theatre — a reviewer clicking through diffs at a rate that proves nobody read them. At L4, no cap, so one bad run touches everything. At L5, no separate policy log, so you can describe what happened but not why it was allowed.

Where This Analysis Is Uncertain

Four things in this piece are weaker than the confident register of the rest, and it is worth being explicit about them.

We have not run this. LeapForce has not conducted a controlled deployment of a vertical agent inside a customer's CRM, EHR or applicant tracking system and measured the outcome. The ladder is a synthesis of deployment-layer patterns and the published evidence cited here, not a finding from our own trial. Where a named study exists we have used it; where one does not, the claim is reasoning, and we have tried to mark the difference.

The evidence base skews to a few verticals. The one randomised trial we could cite is in clinical documentation. Recruiting and lending have clear regulation but little published outcome data. For sales, we have analyst prediction rather than measurement. Anyone claiming to know the production success rate of vertical agents across industries is extrapolating, including us.

The regulatory picture is moving. NIST's work on agent identity and authorisation was still a concept paper at the time of writing, not a control set. EU AI Act obligations for high-risk systems phase in over a period that continues past this article's date. Treat the rules cited here as the floor that already exists, not the final shape.

The ladder is a simplification. Real deployments straddle rungs: an agent may be L1 on one object and L4 on another, and a single "write" can be reversible in the database and irreversible in its downstream effects. The ladder is useful because it forces a specific question, not because six categories exhaust reality. If your workflow does not fit a rung cleanly, that is information about the workflow.

A note on sources. Reddit's practitioner threads on this topic were not reachable through any of our fetch paths for this piece, so the field voices here come from Hacker News, which skews toward platform and infrastructure engineers rather than the recruiters, clinicians and sellers who will actually use these agents. The underlying Gartner research reports are available to clients only; we have cited the public press releases, which contain the figures but not the methodology behind them.

What a Governed Vertical Agent Needs Underneath

Everything above is a governance argument, so it is fair to say what we build. LeapForce is a deployment layer, not a vertical agent: we do not sell a recruiting agent, a clinical scribe or a sales agent, and if you need one of those, buy one of those. What we build is the layer underneath (access, policy, cost and audit) so that whichever vertical agent you choose can be given a rung and held to it.

Four pieces map directly onto the ladder. Access and Identity treats non-human identities as first class: every agent has an owner, a scope and an expiry, and offboarding is one step rather than an archaeology project. Connectors is a registry IT vets once, with action-level scoping, credential brokering, and human-in-the-loop gates — the machinery an L3 rung needs in order to exist at all. Observability and Audit is where the answer to "why was this allowed?" comes from: as that page puts it, the trail "records what was refused, not only what ran". And AI Coworkers exists so a proven workflow becomes a named, owned company asset rather than someone's personal automation, through a lifecycle we run as Build → Scope → Review → Share → Improve; the Scope step is literally where a coworker is given read and write boundaries, and ownership survives the person who built it leaving.

Our gateway rollout model is deliberately staged the same way, and we state it as Observe first. Enforce second. Optimize third. Point one team's traffic at the gateway in observe mode, learn what is actually running, then turn on enforcement, then tune cost. It is the same instinct as the ladder: find out what rung you are really on before you decide what to control.

One more thing, in the spirit of the rest of this piece. LeapForce is in active development, and our own convention is that per-capability build status (live, in development, roadmap) is disclosed honestly on request rather than blurred into a feature grid. If your vertical agent programme needs a specific control in production by a specific quarter, ask us which bucket it is in before you plan around it. That is a better conversation than discovering the answer at implementation, which is the same argument this entire article makes about somebody else's agent.

 FAQ

Frequently asked questions

A vertical AI agent is software that carries out a specific job in one industry from start to finish, inside that industry's own systems, instead of answering questions about the job. A claims agent reads the claim, applies the payer's rules and produces a coded submission. The defining trait is not that it knows the vocabulary — most current models do — but that it holds credentials into the system of record and is permitted to change things there.

Vertical SaaS is an instrument a trained person operates: it shows forms and dashboards, and the human supplies judgement and keystrokes. A vertical AI agent supplies both, and the human supplies the objective and, ideally, the approval. Commercially you shift from buying seats to buying outcomes. Operationally the bigger change is in access control: seat-based permissions assume a person behind every session, and with agents there is not one.

For most enterprise workflows, retrieval over your own documents plus well-specified tools handles the domain gap, and fine-tuning is a cost and staleness liability without a matching benefit. Fine-tuning earns its place for consistent output formats, unusual notation, or latency and cost targets a smaller tuned model can hit. Neither approach gives the agent an identity, a scope, an approval path or a log, which is where the harder work sits regardless.

Determine your rung on the Record Authority Ladder first, then evaluate only what enforces that rung. Four questions do most of the filtering: can the agent hold an identity distinct from both a shared service account and a human user; can permissions be scoped per action rather than per connection; are policy decisions logged separately from execution; and can you revoke everything it reaches in one step. Domain depth is the last filter, because it is now the most common capability on the shortlist.

There is no single credible figure, and any article that gives you one is guessing. The cost has four components that behave differently: subscription or platform fees, model consumption that scales with usage, integration and permissions work that is usually the largest line in year one, and ongoing ownership as processes change. Ask any vendor for the consumption profile of a comparable customer at your volume, and budget the permissions work separately — it is a security project, not an integration task.

Realistically, the timeline is governed by access reviews, not by model work. Getting to L2, a drafting agent whose output stays outside the record, is fast. Often weeks, because it changes nothing. Getting to L3 or L4 depends on how long it takes your organisation to agree what the agent may write, obtain the scoped credentials, and stand up approval and logging. In regulated verticals add whatever your audit obligation requires; under NYC Local Law 144, for example, a bias audit must precede use, not follow it.

Gartner attributes its prediction that over 40% of agentic AI projects will be cancelled by end-2027 to escalating costs, unclear business value, and inadequate risk controls. Our reading is that the third cause is what converts the first two into a cancellation: a pilot that produced good drafts cannot be promoted to unattended execution because nobody can answer who the agent is, what it may write, and how a refusal is proved. The pilot did not fail. It ran out of governance.

If the agent screens job applicants, evaluates creditworthiness, or determines eligibility for essential public services including healthcare, then Annex III classifies it as high-risk and the obligations attach to you as the deployer, not only to the vendor. The classification follows intended use, not the technology, so calling it an assistant does not change it. Our guide for deployers walks through the obligations in more depth.

Configuring one, often yes — most vertical vendors and horizontal platforms have no-code builders, and that is a real advance. Deploying one safely, no. Someone has to decide the scope, obtain credentials that are not a shared account, wire the approval gate, and confirm the log answers audit questions. Those are engineering and security decisions wearing a configuration interface. The no-code claim is honest about the builder and quiet about the permissions.

This is the question that most often turns a successful pilot into an orphan. If the agent runs on an individual's credentials or personal account, it stops or becomes unattributable the day they leave. Assign a named owner who is a role rather than a person, give the agent its own identity with an expiry, and record the ownership somewhere that is reviewed. We have written in more detail about promoting personal workflows into owned company assets.

Measure three things, not one. Task outcome: did the record end up correct, checked by sampling rather than by the agent's own confidence. Human effort: time actually removed, which the ambient scribe trial shows can be zero even for a well-built domain product. And exception rate: how often a human had to intervene, trending over time. If the exception rate is not falling, you have automated the easy cases and moved the hard ones to a worse interface.

Four clauses, in order of how often they are missing. Identity: the agent authenticates as a distinct non-human identity you can enumerate and revoke. Scope: permissions are expressed per action, with a written list of writable fields. Evidence: you can export a log that includes denied actions and the policy that denied them, in a form an auditor accepts. Exit: on termination you retain the decision history, because the obligation to explain a decision outlives the subscription.

Ready to Govern Your AI?

Talk to LeapForce — one controlled layer for every AI tool, connector, model, and agent.

Thirty minutes · No pitch deck

Ready to turn AI experiments into measurable ROI?

Bring one outcome you'd like AI to move. We'll help you scope a pilot you can actually measure — and tell you honestly if it's not worth doing yet.

Comments