What Can AI Do for Business? Only What You Wrote Down

What AI can do for business is bounded by one thing: it can do any job whose inputs are already written down somewhere a machine can read, whose standard for a

What AI can do for business is bounded by one thing: it can do any job whose inputs are already written down somewhere a machine can read, whose standard for a good answer can be shown in examples, and whose result lands in a system that logs it. That is the whole boundary. Every use-case list you have read is a list of jobs that happen to sit inside it.

Our position at Leapforce is that the popular question is badly framed. Lists of AI business use cases answer it with departments — customer service, marketing, finance, recruiting — as if the deciding factor were which team owns the work. It is not. Two jobs in the same department, of identical difficulty, land on opposite sides of the line depending on whether the knowledge the job needs exists in text or only in the head of the person who has done it for nine years. Sort by that and the menu turns into a map.

The gap shows up the moment someone tries to run the project. In June 2026 a developer posting as iExploder opened an Ask HN thread on how AI adoption is driven inside companies, describing executive mandates that feel like "being told what compiler/IDE to use by people who have no idea about engineering" (HN item 48518427). Further down, a commenter as inthepond, who does this work for small and large firms, named the actual blocker in one clause: "It's always difficult for business to clearly state their SOP from 0 to 1" (HN item 48536231). The executive is asking what AI can do. The practitioner is saying nobody can describe the work.

The short answer: AI can do the parts of your business that are already written down — conversations you keep, documents that arrive in a fixed shape, systems of record with typed fields, code, and content you have published — and it reliably disappoints everywhere the deciding knowledge lives in someone's memory, no matter which model you buy.

Last updated: July 30, 2026.

Five zones where a machine-readable record already exists, and the business work that sits outside that boundary

The written-down boundary. Work moves left only when someone creates the record.

What the Question Is Really Asking

When an owner or an operations lead asks what AI can do for their business, they are not asking for a taxonomy. They are asking a scope question: how much of what happens here is in play, and how would I recognise the parts that are? A department list cannot answer that, because it describes where work sits rather than what the work is made of.

Here is the substitution that makes the question answerable. Every workflow in a company consumes some input, applies some standard, and produces some result. A language model can participate only where all three of those exist in a form it can be handed. It has no access to your building, your customers' tone of voice on the phone last Tuesday, or the reason your best account manager stopped calling one client in April. It has access to text and structured data. That is not a limitation you can buy your way out of with a larger model, because it is not a reasoning limitation at all.

This also explains an oddity that catches people out constantly. Two jobs described in the same sentence can behave completely differently. "Draft the customer reply" works, because five years of resolved tickets show what a good reply looks like. "Draft the price concession" does not, because the reasoning behind past concessions was verbal, the exceptions were personal, and the only written trace is the final number with no record of why.

So the useful reframe for anyone evaluating AI for business is this: stop asking which tasks AI is good at, and start asking which of your workflows leave a paper trail. The rest of this piece turns that into something you can run on Monday, then gives the five places where the paper trail already exists whether you planned it or not.

What AI Is Actually Doing in Companies Right Now

Adoption is broad and shallow. According to the 2026 AI Index Report from Stanford HAI, organizational AI adoption rose to 88% of surveyed organizations in 2025, with generative AI used in at least one business function at 70% — yet the same chapter notes that "AI agent deployment was in the single digits across nearly all business functions." Almost everyone has AI somewhere. Almost nobody has it acting on their behalf. That gap is most of the honest answer to what AI can do for business right now.

The picture from firm-level government data is soberer still and worth holding against the survey numbers, because the two measure different things. A Federal Reserve FEDS note by Jeffrey S. Allen, published 3 April 2026, puts firm-weighted AI adoption at 18% by the end of 2025 while noting that 78% of the labor force works at an AI-adopting firm. Most firms have not adopted; most employees work somewhere that has. Sector spread is wide in exactly the direction this article predicts: professional services 33% and financial services 30% against manufacturing around 15%, wholesale trade 13%, and accommodation and food services 8%. The high-adoption sectors are the ones whose product is already documents and records.

Where value has been measured, it clusters the same way. The AI Index summarises productivity gains as roughly 14 to 15% in customer support, 26% in software development, and a 50% lift in marketing output. All three are functions whose raw material is text that a company already stores.

SignalFigureSource
Organizations using AI in at least one function88% of surveyed organizations, 2025Stanford HAI, 2026 AI Index
Generative AI in at least one business function70% of organizationsStanford HAI, 2026 AI Index
AI agent deployment by business functionsingle digits across nearly all functionsStanford HAI, 2026 AI Index
US firms that had adopted AI, firm-weighted18% by end of 2025Federal Reserve FEDS note, Apr 2026
Labor force at AI-adopting firms78%Federal Reserve FEDS note, Apr 2026
Professional services / finance / manufacturing / wholesale / food service adoption33% / 30% / ~15% / 13% / 8%Federal Reserve FEDS note, Apr 2026
Measured productivity gain, customer support14 to 15%Stanford HAI, 2026 AI Index
Measured productivity gain, software development26%Stanford HAI, 2026 AI Index

One recognisable source is missing from that table and the absence is worth naming. The US Census Bureau's Business Trends and Outlook Survey is the underlying instrument for the Federal Reserve figures and publishes its own AI-use estimates, but census.gov blocked every fetch we attempted, including a browser session. We are therefore citing the Federal Reserve's analysis of that survey rather than the survey pages directly, and we have not independently reproduced the widely quoted 2026 BTOS headline rate.

Why the shallow part matters more than the broad part

A department list implies the constraint is imagination: you simply have not thought of the thirteenth use case yet. The measured record says otherwise. Adoption is near-universal and autonomous deployment is in the single digits, which means the hard part sits after someone has already picked a use case. That is consistent with what practitioners report: the blocker is not choosing, it is that the chosen job turns out to need knowledge nobody ever typed.

The Written-Down Test

The Written-Down Test is three questions asked in order about one candidate job. It takes about ten minutes per candidate and it replaces the guesswork of picking from a menu. Each gate asks for a record, meaning a durable and readable artifact, not a person's willingness to explain.

Gate 1, the input record. Is everything the work needs already captured somewhere a system can read? Not summarised in a manager's head, not in a call nobody recorded. If the answer is "mostly, except the bit where we check with Dave", the honest answer is no, and Dave's bit is the project.

Gate 2, the standard record. Can you produce twenty examples of a good result and say, per example, why it is good? This is the gate people skip, and it is the one that decides whether you can ever measure the thing. A model can be pointed at any job; it can only be graded on jobs where "correct" is demonstrable. If good is a feeling in a senior person's chest, you have no grading key and no way to know when quality slips.

Gate 3, the outcome record. Will the result land somewhere you can count a month later? A drafted reply that gets sent from a mailbox with no logging produces no evidence. Absence of an outcome record is why so many pilots end with everyone agreeing it felt faster and nobody able to fund round two.

Three ordered gates of the Written-Down Test, with what to do on a yes and on a no

Run the three gates before you shortlist a tool, not after.

The order matters. A missing input record makes gates 2 and 3 unanswerable, so stop at the first no rather than scoring all three. And the no is useful output, not a rejection: it tells you precisely which record to create, which is a scoping decision you can hand to a person.

There is a reason we trust this gate ordering rather than a difficulty ranking, and it comes from two studies that appear to contradict each other until you look at what was written down in each case. The NBER working paper Generative AI at Work by Erik Brynjolfsson, Danielle Li and Lindsey R. Raymond studied 5,179 customer support agents and found issues resolved per hour rose 14% on average, including a 34% improvement for novice and low-skilled workers, with minimal impact on experienced and highly skilled ones. The authors describe the mechanism directly: the model "disseminates the best practices of more able workers." Those best practices existed in text, in resolved tickets, which is exactly why they could be disseminated.

Now the other direction. METR's randomized controlled trial, published 10 July 2025, gave 16 experienced open-source developers 246 real issues in repositories they knew well. With AI tools allowed, they took 19% longer. They had expected a 24% speedup and, after the fact, still believed they had been sped up by 20%. Among the candidate explanations METR offers is that AI performs comparatively poorly in settings "with many implicit requirements (e.g. relating to documentation, testing coverage, or linting/formatting) that take humans substantial time to learn."

Read together: the model won where the standard was in the archive, and lost where the standard was in the practitioner's head. Same technology, opposite results, and the discriminator is legibility rather than task difficulty. Note the METR caveat honestly. Sixteen developers is a small sample, the tools tested date from early 2025, and the authors themselves list several candidate explanations rather than one. It is a data point about a mechanism, not a law about coding assistants.

The Five Zones at a Glance

Five kinds of record already exist in almost every company, usually as a byproduct of something else. They are where the answer to "what can AI do for us" actually lives, and they are ordered here by how little work you have to do before starting.

ZoneWhy the record existsWhat AI does here todayRealistic first projectMain failure mode
1. Conversations you keepthe conversation is the recordsummarise, tag, route, draft repliesdraft support replies for human sendreviewer becomes the bottleneck
2. Documents in a fixed shapea counterparty sends themextract to fields, compare to terms, flag exceptionsinvoice header and line extractionexception handling costs more than the extraction saves
3. Systems of record with typed fieldssoftware forced the fieldquery in plain language, reconcile, propose writesanomaly and stale-record reportingfields nobody fills produce confident nonsense
4. Code, config, infrastructureit must be machine-readable to runwrite, review, test, explain, documenttest and documentation generationimplicit house standards nobody wrote down
5. Content you have publishedyou published itrepurpose, vary, translate, localiseturning long assets into channel variantsoutput volume mistaken for outcome

Each zone below follows the same four-part shape so you can compare them: what the record is, what AI does there today, what it cannot do there, and how you would know it worked.

Zone 1: Conversations You Already Keep

The record here is free. Support tickets, email threads, meeting transcripts and chat history exist because the conversation and the record are the same object, which makes this the cheapest zone in any company and the reason support is the most-cited AI success story in the field.

What AI does here today. Compress a long thread into a decision summary. Tag and route incoming customer service messages by intent. Retrieve the three prior cases that resemble this one. Draft a reply grounded in those cases for a person to send. Extract commitments and owners from a meeting transcript. All of it is text-in, text-out, with a human reading the output as part of the existing job.

What it cannot do here. It cannot know what was agreed on a call nobody recorded, and it cannot represent a customer relationship that lives in the account manager's judgement. It also cannot reliably answer questions whose answers were never in the archive; if your team resolved a class of problem by walking to someone's desk, the tickets show the symptom and never the fix.

How you would know it worked. Resolution time and reply time per ticket, sampled by segment rather than averaged, plus an error-rate sample of drafted replies graded against your gate-2 examples. The Brynjolfsson study's heterogeneity is the thing to watch for: gains concentrated in less experienced staff and near-zero for your best people, which changes who you deploy it to and how you talk about it internally.

The trap. A drafting assistant moves work from writing to reviewing. If reviewing a draft takes 80% as long as writing from scratch, you have bought very little, and you will not find that out from a demo. Time the review step during the pilot, with the actual reviewers, on a normal week.

Zone 2: Documents That Arrive in a Fixed Shape

The record arrives from outside, in a form somebody else controls. Invoices, purchase orders, claims, remittance advices, contracts, application forms and CVs all share a property that makes them tractable: the fields you want are named on the page, and the same fields appear on every instance.

What AI does here today. Turn an unstructured or semi-structured document into typed fields. Compare a received document against agreed terms and flag the difference. Sort a stack into "matches, needs a look, is wrong." Answer questions across a contract set, with the clause cited. This zone is where language models genuinely outperform the previous decade of template-based extraction, because they cope with layout variation that broke rules.

What it cannot do here. It cannot decide what to do about the exception. It cannot resolve a disagreement about what a clause means, and it should not be the last thing that looks at a document before money moves. It also cannot invent a field the sender did not include, which sounds obvious and is the source of most silent errors: a missing purchase-order number becomes a plausible guess unless the pipeline is built to refuse.

How you would know it worked. Field-level accuracy on a held-back sample you label by hand, per field rather than per document, because the average hides the one field that matters. Then the exception rate, and the minutes a human spends per exception. Multiply those two before you believe any business case.

The arithmetic that decides this zone. Extraction at 95% per document sounds excellent until you notice that a 5% exception rate on 4,000 documents a month is 200 human interventions, and that an intervention on a document a person did not read from scratch is often slower than the original task. This is why document projects succeed at high volume and fail at medium volume: the fixed cost of building the exception path is the same either way.

Zone 3: Systems of Record With Typed Fields

Here the record exists because software refused to accept anything else. Your CRM, ledger, ticketing system, HRIS and inventory tables contain clean, typed, queryable rows, and that is a materially different asset from a pile of documents.

What AI does here today. Translate a plain-language question into a query and explain the result. Reconcile two systems and report the differences. Spot rows that look wrong against a stated rule. Summarise a pipeline or a backlog into something readable at a management level. Propose writes: an updated stage, a corrected field, a created record. which a person confirms.

What it cannot do here. It cannot fix field discipline. If your reps do not fill the close-reason field, no model will tell you why deals are lost; it will produce a fluent narrative from the fields that are populated, which is worse than silence because it sounds like analysis. And the moment you allow writes rather than proposals, you have crossed from a reporting project into an access-control project, which is a different kind of decision with a different owner.

How you would know it worked. Whether a specific recurring question is answered without an analyst in the loop, timed. For proposed writes, the acceptance rate of proposals and the count of accepted-then-reverted changes, which is the honest error signal.

The write line. This is the boundary where "what can AI do for us" quietly becomes "what are we letting it do." We have written the full version of that argument in our earlier analysis of AI agent use cases ranked by blast radius. The short form is that read and draft are Monday-morning projects, and write is a governance project first.

Zone 4: Code, Configuration and Infrastructure

Software is the most thoroughly written-down part of most companies, because it has to be machine-readable to run at all. Repositories, schemas, pipeline definitions, tests and runbooks are a complete specification of behaviour in text, which is why this zone shows the largest measured gains and the loudest disagreements.

What AI does here today. Write and modify code against a described intent. Review a change and name likely defects. Generate tests and fixtures. Explain an unfamiliar system to someone who has to change it. Draft the documentation and the migration notes that nobody wanted to write. The AI Index puts software development gains at 26%, the highest of the functions it summarises.

What it cannot do here. It cannot infer the standards you never wrote down, which is precisely the METR finding above: experienced developers on mature codebases with many implicit requirements went 19% slower. It cannot know which of three correct implementations your team will accept in review. And it cannot own the consequence of a change, which is why the reviewing engineer's name still has to be on it.

How you would know it worked. Not lines produced. Cycle time from first commit to merged, review iterations per change, and change failure rate, compared against the same team's baseline before the tool. If review iterations climb while cycle time falls, you have moved effort rather than removed it.

The uncomfortable part. METR's developers believed they had been sped up 20% while being slowed 19%. That gap is the strongest argument in this entire article for gate 3. Self-reported time savings are not evidence, and a zone this well documented is exactly where an organisation is most likely to mistake activity for progress.

Zone 5: Everything You Have Already Published

Your website copy, product documentation, past campaigns, decks, help centre and knowledge base are a corpus you own outright, with known provenance and no privacy question attached. That combination makes this the zone with the fewest governance obstacles and the least defensible business case.

What AI does here today. Turn one long asset into channel variants. Translate and localise. Draft first versions in an established voice. Answer customer questions from documentation with the source cited. Rewrite for a different audience or reading level. The AI Index reports a 50% lift in marketing output, and the word doing the work in that sentence is output.

What it cannot do here. It cannot tell you whether more output helps. It cannot originate a position, because a position is a commitment about the world and the corpus only contains commitments you already made. It cannot fix a message that was not landing; it will produce more of it, faster, in more formats.

How you would know it worked. Pipeline or conversion attributable to the new variants, not volume published. If the only metric that moved is the count of assets, the honest verdict is that you automated a cost centre and left the outcome untouched.

Why we rank this last despite the easy record. Every other zone removes work someone was already doing. This one mostly increases supply, and increased supply of undifferentiated content is now competing with everyone else's increased supply. Do it, but do not lead your AI programme with it.

The Unwritten Zone: Where AI Reliably Disappoints

Outside the boundary sits a body of work that resists AI not because it is difficult but because the deciding knowledge was never recorded. This is the section most use-case lists omit, and it is the one that determines whether your programme reads as honest inside the company.

Pricing and negotiation judgement. Your systems hold what you charged. They do not hold why you conceded, which competitor was in the room, or that the discount was a favour to a person who has since left. Fed the outcomes without the reasoning, a model learns the wrong rule and states it confidently.

Relationship context. Which client tolerates a chase in week one and which will churn over it. Who at the account actually decides. Why nobody has emailed that contact since the reorganisation. Occasionally a fragment shows up in CRM notes; the operative part almost never does.

Unwritten quality standards. The "we would never ship it like that" reflex. This is METR's implicit-requirements finding restated in business terms, and it is the most expensive item on this list because it is invisible until the output is rejected. It is also the most fixable: writing the standard down is a real, scoped piece of work with a real deliverable.

Why the exception exists. One customer is invoiced differently, one region skips a step, one product ships without the usual check. The rule is in the system. The reason is in a conversation from 2019. A model applying the rule without the reason will generalise it to cases where it does not hold.

Institutional memory held by one person. The operator who knows which supplier slips in August. When that knowledge leaves, it leaves whether or not you deployed AI, which is an argument for capturing it independent of any tool decision.

The pattern across all five: the failure is not that the model is not smart enough. It is that the input was never created. This is the same underlying observation the MIT Project NANDA report reached from a different direction. Its widely quoted claim that 95% of generative AI pilots delivered no measurable P&L impact is attributed to tools that, as summarised by the AI Governance Library, "fail not because of poor models, but because they don't learn, adapt, or integrate." Treat the 95% with care: the study rests on 300-plus initiative reviews, 52 interviews and 153 survey responses, its authors do not claim it is representative, and its success bar is a marked and sustained impact. We cite it for the mechanism, not the headline number.

Two Jobs That Need a Different Technology

Two items appear on nearly every "what can AI do for business" list and belong in a separate category, because reaching for a language model is the wrong move even when the record exists.

Demand and inventory forecasting. This is a numeric time-series problem, and it has been solvable with statistical and machine-learning methods for years, well before language models existed. What it needs is history: enough periods, enough SKUs, and clean records of stockouts and promotions. A chat interface over your inventory table is not forecasting; it is querying. If a vendor presents a language model as your demand planner, ask what it is trained on and what the backtest looks like. If a small business has two years of noisy sales history, the honest answer is that no technology on this list will produce a reliable forecast from it, and the useful project is a reorder rule with a safety margin.

Decisions about people. Screening candidates and assessing creditworthiness are the two most common list items that carry legal weight in the EU. Annex III of the EU AI Act classifies as high-risk "AI systems intended to be used for the recruitment or selection of natural persons, in particular to place targeted job advertisements, to analyse and filter job applications, and to evaluate candidates", and separately systems "intended to be used to evaluate the creditworthiness of natural persons or establish their credit score, with the exception of AI systems used for the purpose of detecting financial fraud." High-risk does not mean forbidden. It means a documented obligation set: risk management, data governance, logging, human oversight. That turns a two-week automation into a compliance programme. Our guide to EU AI Act compliance for deployers covers what falls on you as a deployer rather than a provider.

Job commonly listedWhat it actually needsRight first move
Inventory and demand forecastingnumeric history, enough periods and volumestatistical forecasting or a reorder rule, not a chat model
Candidate screeningEU AI Act high-risk obligationskeep humans deciding; AI summarises, never ranks
Credit and risk scoringhigh-risk obligations plus model governancetreat as a regulated model, not an AI feature
Financial close and reportingtyped ledger fields and reconciliation rulesZone 3 anomaly reporting, human sign-off

What It Costs to Write Something Down

The uncomfortable implication of the test is that when a job fails a gate, the project is not an AI project. It is a documentation project, and somebody has to be paid to do it. This is the cost that no use-case list prices, and it is the difference between a programme that compounds and one that produces a series of demos.

Three things are being bought, and it helps to name them separately.

Capturing the input. Recording what was previously verbal: turning the call into a logged note, the walk-to-the-desk fix into a ticket resolution, the spreadsheet on a laptop into a table with a schema. The cost is mostly behavioural, which is why it is underestimated. A field that people do not fill is not a record.

Writing the standard. Producing the twenty graded examples gate 2 asks for. This is genuinely skilled work, it takes your most experienced person's time, and it is the most reusable document in the whole exercise, because it doubles as the acceptance test, the training material and the review rubric.

Building the outcome log. Making sure the result lands somewhere countable. Usually the cheapest of the three and the one most often skipped.

There is survey evidence that this is where organisations are stuck, though it comes from a vendor and should be read as such. Cloudera's Data Readiness Index, published 14 April 2026 and based on 1,270 IT leaders at companies with 1,000-plus employees, reports that nearly four in five say their AI and data initiatives are constrained by limited data access across environments. A storage vendor has an obvious interest in that finding. It is still consistent with the practitioner account from the Ask HN thread, with the METR mechanism, and with the sector pattern in the Federal Reserve data.

We will not put a dollar figure on the documentation work, because it varies with how much of your process is verbal and we have not run a controlled measurement of it. What we will commit to is the shape: budget the record before the licence, and expect the record to be the larger line item on your first use case and a much smaller one on your fourth, because zones share records. The support archive that powers reply drafting also powers routing and knowledge retrieval.

The Test, Run on Three Real Candidates

Here is the test applied to three jobs that appear on nearly every list, using the same three gates, so the scoring is comparable rather than impressionistic.

CandidateGate 1: input recordGate 2: standard recordGate 3: outcome recordVerdict
Draft replies to inbound support emailyes — five years of resolved threadsyes — resolved tickets are graded examplesyes — helpdesk logs send time and reopen ratebuild it, human sends
Set reorder quantities for 900 SKUspartial — sales history yes, stockouts and promos often not loggedno — "right quantity" was judgement, never writtenyes — stockouts are countablefix the stockout log first; use a rule, not a model
Produce the monthly board summarypartial — the numbers are in systems, the narrative is verbalno — nobody has written what a good summary containsweak — nobody grades itwrite the standard first; three past summaries annotated

Two of three fail, and both failures point at a specific missing document rather than at a missing tool. That is the test earning its keep. The reorder case fails gate 2 in a way no model release will change, and the board summary fails on a document one person could write in an afternoon. That makes it, counterintuitively, the fastest of the three to unblock.

Worth noticing: the candidate that passes is the one the field has the most evidence for, and it passes for the reason the evidence exists. Support archives are the most thoroughly written-down process in most companies.

Widening the Boundary on Purpose

Once you accept that the boundary is made of records, expanding what AI can do for your business becomes a deliberate programme rather than a waiting game. You are not waiting for better models. You are choosing which record to create next.

The ordering rule we use: create the record that unlocks the most zones, not the record for the most exciting use case. Records are shared infrastructure and use cases are not.

Record you createCheapest way to create itWhat it unlocks
Resolution notes on every closed ticketmake the field required, sample it weekly for qualityreply drafting, routing, self-service answers, gate-2 examples for free
Meeting decisions and owners, in textone recorded and transcribed meeting type, not all of themcommitment tracking, handover summaries, project status
A written quality standard per output typeyour most experienced person, twenty annotated examplesgrading, review rubrics, onboarding, every gate-2 answer
Close-reason and loss-reason on every dealrequired picklist plus one free-text linepipeline analysis that is not fiction
A schema for the spreadsheet everyone actually usesmove it into a table with typesZone 3 querying and reconciliation

Two cautions on sequencing. First, do not create records you will not use; a documentation programme with no consuming use case dies within a quarter, so pair each record with the one job it unblocks. Second, make the record a byproduct of work people already do rather than an extra step, because the failure mode here is not refusal but quiet decay. The field gets filled with "n/a" and you now have a record that lies.

If you want the sequencing question treated properly, meaning which order to run the jobs in once several have passed, our analysis of why AI pilots stall before production covers the ground this article deliberately does not.

What Changes the Moment AI Touches a Real Record

Everything above is about scope. There is one consequence of widening scope that belongs in this article because it arrives immediately and surprises people: the moment a model reads a real company record, you have granted access to that record, and the grant now needs an owner, a boundary and a log.

That is the layer Leapforce builds. One controlled path to every model, with the caller identified, the policy checked before the request leaves, and the call recorded against a person, team and budget. Our published rollout model for it is deliberately unglamorous — observe first, enforce second, optimize third — because you cannot write sensible policy for AI use you have not yet measured, and phase one is just pointing one team's traffic at a governed gateway in observe mode to find out what is actually happening. The connector side matters for the same reason: when a job moves from reading a record to writing one, the useful control is action-level scoping on the connector rather than a policy in a wiki. One honest disclaimer: Leapforce does not write your standard record or your SOP. That part is your people, and no governance layer substitutes for it. Some capabilities on our platform are still in development and the site labels build status per capability.

Limits of This Analysis

The Written-Down Test is a scoping heuristic, not a proof, and there are places where it will mislead you.

It says nothing about value. A job can pass all three gates and be worth almost nothing, which is why the test filters candidates rather than ranking them. Pair it with a value estimate before committing.

It undersells retrieval and synthesis on messy corpora. Modern systems do extract usable signal from archives that no human had organised, so "the record is a mess" is not the same as "there is no record." The test asks whether a record exists, not whether it is tidy.

The boundary moves. Multimodal models read images and video, so a photographed form or a recorded call can become a record without anyone typing. Anything we say about the boundary's position has a shelf life. The mechanism, that a model can only work from what has been captured, is the durable part.

We did not run a first-hand test for this article. Nobody on our side ran the Written-Down Test across a sample of companies and measured how often each gate fails, so the three worked candidates above are illustrative reasoning from public evidence rather than measurements, and the ordering of the five zones is a judgement, not a ranked result.

The evidence base skews. The strongest measured findings available are in support and software engineering, because those are the functions researchers can instrument. There is much less rigorous public evidence for finance, operations and HR, and the sector adoption gaps in the Federal Reserve data suggest the picture there may be genuinely different rather than merely unmeasured. The Ask HN voices are also a skewed sample, a technical forum speaking about company-wide adoption — and census.gov being unreachable at all three fetch tiers means the firm-level picture reaches you through the Federal Reserve's reading of it rather than ours.

 FAQ

Frequently asked questions

What AI can do for business at small scale is the same five zones, and the cheapest starting points are usually reply drafting from an existing email or ticket archive, extracting fields from supplier documents, and repurposing content you have already published. What changes with size is volume: document extraction needs enough monthly documents to justify building the exception path, so a small business is often better served by drafting and summarising than by extraction. Government data supports the size gap. The Federal Reserve's April 2026 note puts firm-weighted adoption at 18% by end of 2025, with the smallest firms lowest.

Run the Written-Down Test on three candidates and start with the one that clears all three gates rather than the one with the biggest claimed prize. If none clears, the winner is the candidate whose failing gate is cheapest to fix. Usually that is gate 2, because writing twenty graded examples takes one experienced person an afternoon and is reusable everywhere. Do not shortlist tools before a candidate has passed; the tool is the last decision, not the first.

It cannot use knowledge that was never captured. Pricing reasoning, relationship context, unwritten quality standards and the reason an exception exists are all invisible to it, and a larger model does not help because the input does not exist. It also cannot own a consequence, because accountability stays with a named person. And it should not be the last thing that checks a document before money moves.

Not for zones 1, 2 and 5, where mainstream tools cover summarising, drafting, extraction and content work without engineering. You will need engineering as soon as the work touches a system of record, because connecting to a CRM or ledger, handling exceptions and writing back safely are software problems rather than prompt problems. The rough line: reading and drafting rarely needs code, writing to a system always does.

Budget three items, not one: the licence or usage cost, the integration and exception-handling build, and the documentation work the Written-Down Test exposes. On a first use case the documentation is often the largest of the three, and it is the one no vendor quote includes. We are not going to publish a dollar range we have not measured; what we will commit to is that the record cost falls sharply on your second and third use case, because zones share records.

There is no universal threshold — the number that matters is the cost of a wrong answer times the rate of wrong answers, minus what a reviewer catches. A 95% accurate extraction over 4,000 documents a month is 200 exceptions to handle by hand, which can easily exceed the saving. Measure accuracy per field on a hand-labelled held-back sample, then measure how long a human takes on an exception, and only then decide.

It is a governance decision rather than a yes or no. The requirements are the same ones you would apply to any system reading that data: an identified caller, a scoped grant, policy evaluated before the data leaves your boundary, and a log of what was accessed. The failure mode in practice is not a hostile model, it is staff pasting customer data into personal accounts because the approved path is slower, which is why our analysis of shadow AI treats an unusable approved route as the actual risk.

On the current measured evidence it mostly assists, and unevenly. The NBER study of 5,179 support agents found a 14% average gain in issues resolved per hour, but a 34% gain for novices and minimal impact on experienced staff. That is a compression of the skill gap rather than a removal of the role. Stanford's 2026 AI Index also reports AI agent deployment in the single digits across nearly all business functions, which is not what a replacement wave looks like.

If a candidate clears all three gates, a read-only or draft-only version is typically a matter of weeks, because nothing needs to be created first. If it fails a gate, the honest timeline is however long the missing record takes, plus those weeks. The projects that overrun are almost always the ones that started before the record existed and discovered it during the build.

The most cited figure is the MIT Project NANDA claim that 95% of generative AI pilots delivered no measurable P&L impact, and the stated mechanism is tools that do not learn, adapt or integrate rather than weak models. Treat the number cautiously: it rests on 300-plus initiative reviews, 52 interviews and 153 survey responses, with a demanding definition of success. The mechanism is the useful part, and it matches gate 3. A pilot with no outcome record cannot prove anything and therefore cannot be funded again.

Annex III of the EU AI Act lists eight areas, and two show up constantly on business use-case lists: recruitment and candidate evaluation, and evaluating the creditworthiness of natural persons. Both bring documented obligations covering risk management, data governance, logging and human oversight. High-risk is not banned, but it converts a short automation project into a compliance programme, so do not let it be your first use case.

Measure the outcome the business already tracked before the tool existed, sampled rather than averaged, against the same team's pre-tool baseline. In support that is resolution time and reopen rate; in engineering, cycle time, review iterations and change failure rate; in document work, per-field accuracy and minutes per exception. Do not measure self-reported time saved: METR's developers believed AI had sped them up 20% while a controlled measurement showed them 19% slower.

A named person, with the same clarity you would apply to a system rather than a tool. Once something reads or writes company records on a schedule, it needs an owner who can answer what it can reach, who approved that, and what happens when the owner leaves. That is the argument we set out in our earlier analysis of giving an AI agent an owner, a scope and an expiry. The most common governance gap we see described is an automation still running under the personal credentials of someone who has moved teams.

Ready to Govern Your AI?

Talk to LeapForce — one controlled layer for every AI tool, connector, model, and agent.

Thirty minutes · No pitch deck

Ready to turn AI experiments into measurable ROI?

Bring one outcome you'd like AI to move. We'll help you scope a pilot you can actually measure — and tell you honestly if it's not worth doing yet.

Comments