AI for Sales: The Four Sentences an Agent Must Not Send

An AI sales agent can safely do research, summarise, schedule, log, draft and route. It cannot safely make a promise. Four kinds of sentence bind your company w

An AI sales agent can safely do research, summarise, schedule, log, draft and route. It cannot safely make a promise. Four kinds of sentence bind your company when they leave the building: a price, a contract term, a delivery date, and a claim about what the product does. Everything else in AI for sales is a productivity question. Those four are a liability question, and they need a named human owner before send.

That is our position, and it comes from a specific observation: the industry has built approval as a per-campaign toggle when the thing that needs approving is a sentence. You do not want a rep reviewing 400 AI-drafted emails. You want only the sentences inside those emails that create an obligation to be the ones that stop at a desk.

Nothing below is legal advice. We cite settled statute and one published tribunal decision, we say plainly where the law is contested or untested, and every prescriptive artefact in this article is a starting shape for a conversation with your own counsel and deal desk, not a substitute for one.

The problem is being voiced by the people deploying this. In an Ask HN thread on 20 April 2026, a poster writing as u/aegisproxy described shipping agents that read files, call APIs and write to databases, and said the conversation around controlling them is "almost nonexistent". They named the Air Canada chatbot case as their evidence. They are also building in this space, which we note rather than hide; the observation still lands, because we could not find a single AI sales product page that documents a per-sentence approval gate.

The short answer: Let an AI sales agent send anything that only reports a fact you can already prove; require a named human owner to approve any sentence that states a price, a contract term, a delivery commitment, or a product capability, because those four are the ones a court, a regulator, or a customer will hold you to.

Last updated: July 31, 2026.

Four lanes of AI sales output: one free lane for retrievable facts, four gated lanes for price, terms, timing and capability claims

AI for Sales, Defined by What the Company Has to Honour

AI for sales is the use of language models and agents to do the reading, writing, research and record-keeping around a deal. The useful definition for anyone who has to sign off on it is narrower: AI sales automation is the moment your company starts producing customer-facing sentences that no employee wrote. Everything a governance review needs to decide follows from that one fact, because a sentence you did not write is still a sentence you have to stand behind.

Most coverage of AI in sales sorts the work by stage — prospecting, qualification, outreach, discovery, proposal, close, renewal. That is a useful map for buying software and a poor map for deciding what an agent may do, because the risk is not evenly spread inside a stage. Two emails sent in the same stage, by the same agent, on the same day, can differ by three orders of magnitude in what they cost you. One says "here is the case study you asked for." The other says "we can do 22% off for a two-year term." The stage is identical. The exposure is not.

So we sort by a different axis: what does this sentence obligate us to do? We have written before about ranking sales work by how expensive a bad automation is to reverse, which is the right lens when you are choosing which pipeline steps to automate at all. This piece sits one level down, inside the steps you have already decided to automate, and asks a narrower question: of the text this agent produces, which sentences can leave without a human name attached?

Three terms, used consistently below:

TermMeaning here
CommitmentA sentence that, if the recipient acts on it, the company is expected to honour
Commitment classOne of the four categories a commitment falls into: price, terms, timing, claim
Commitment authorityThe named human, role or policy object that is allowed to approve a given class for a given agent

That vocabulary is the whole framework, and it is deliberately small. A control a sales manager cannot hold in their head on a Tuesday will be switched off by Thursday.

What a Tribunal Already Decided About a Company's Bot

The most-cited decision on this question is Moffatt v. Air Canada, 2024 BCCRT 149 (PDF of the reasons for decision; CanLII's copy blocks automated retrieval), issued by British Columbia's Civil Resolution Tribunal on 14 February 2024 (file SC-2023-005609, Tribunal Member Christopher C. Rivers). Jake Moffatt booked a last-minute flight after the death of their grandmother and relied on the airline's website chatbot, which the decision records at paragraph 15 as telling a customer who needed to travel immediately "or have already travelled" that they could submit a ticket for a reduced bereavement rate "within 90 days of the date your ticket was issued." The airline's actual published policy, on the page the chatbot linked to, said the bereavement policy did not apply to requests made after travel had been completed.

Air Canada's defence is the part everyone in sales operations should read. The tribunal recorded that the airline argued it could not be liable for information provided by its own agents, servants or representatives, including a chatbot, and did not explain why. The member's characterisation was blunt: the airline was, in effect, suggesting the chatbot was "a separate legal entity" answerable for its own actions, and the member called that "a remarkable submission." The decision then states the operative principle at paragraph 27: "It should be obvious to Air Canada that it is responsible for all the information on its website. It makes no difference whether the information comes from a static page or a chatbot."

The tribunal found negligent misrepresentation and awarded CAD $650.88 — the difference between the CAD $1,630.36 Moffatt paid and the CAD $979.48 they would have paid at bereavement rates. Six hundred and fifty dollars. That is the entire financial consequence, and it is why the case matters more as a statement of principle than as a risk number. The cost of the ruling was trivial; the cost of the position the airline took in public was not.

Two things this case does not do, which the internet routinely claims it does. It does not create a general rule that every chatbot utterance forms a binding contract; the finding was in tort, for negligent misrepresentation, not in contract. And it is a small-claims tribunal decision in one Canadian province, so it binds nobody outside that context. What it supplies is a quotable rejection of the defence every company instinctively reaches for: the model said it, not us.

The field's reaction is instructive. In the Hacker News thread on the ruling, one commenter noted that where support bots used to quote a real solution or escalate, "now it makes up a non-working solution" (u/WirelessGigabit). In an earlier thread about a dealership whose site chatbot was talked into agreeing to sell a vehicle for a dollar, another wrote that "any screwups are on them" (u/ElevenLathe). Both instincts match where the law already sits. Neither is where most AI sales deployments are configured.

What the Statutes Already Say, and What They Do Not

This is the part to read slowly, because there is a large amount of confident nonsense written about it and the settled material is both narrower and more useful than the speculation.

Automated action is attributable. The federal E-SIGN Act, at 15 U.S.C. § 7001(h), provides that a contract or record "may not be denied legal effect, validity, or enforceability solely because its formation, creation, or delivery involved the action of one or more electronic agents so long as the action of any such electronic agent is legally attributable to the person to be bound." Read the tail of that sentence twice. The statute does not decide attribution for you; it says that once the action is attributable, automation is not a defence.

No human needs to have read it. The Uniform Electronic Transactions Act § 14, as enacted in most US states — California's version is Civil Code § 1633.14 — states that in an automated transaction "a contract may be formed by the interaction of electronic agents of the parties, even if no individual was aware of or reviewed the electronic agents' actions or the resulting terms and agreements." This is the single most load-bearing sentence for anyone deploying AI in sales. The absence of human review is not a gap in formation. It is expressly contemplated.

A promise about the product is a warranty, with no magic words. UCC § 2-313 provides that "any affirmation of fact or promise made by the seller to the buyer which relates to the goods and becomes part of the basis of the bargain creates an express warranty," and § 2-313(2) adds that it is not necessary for the seller to use formal words such as "warrant" or "guarantee," or to have any specific intention to make a warranty. The same subsection carves out puffery: a statement of value, or one purporting to be merely the seller's opinion or commendation, does not create a warranty. That carve-out is the reason "our platform is excellent" is safe and "our platform processes 10,000 invoices an hour" is not.

Deception is deception regardless of who typed it. Section 5 of the FTC Act, 15 U.S.C. § 45(a)(1), declares that "... unfair or deceptive acts or practices in or affecting commerce, are hereby declared unlawful." When the FTC announced Operation AI Comply on 25 September 2024, then-Chair Lina M. Khan put it as plainly as an enforcement agency ever does: "Using AI tools to trick, mislead, or defraud people is illegal," and there is "no AI exemption from the laws on the books."

Now the honest part. There is a real and contested question underneath all of this: whether a specific AI agent's specific statement is attributable to the company in the way E-SIGN requires, and how classical agency doctrine — actual authority, apparent authority, ratification — applies to a system that is not a person and has no mind to hold an intention. We are not going to resolve that for you, and you should distrust any vendor blog that does. The doctrine is being argued, the fact patterns are thin, and the answer will differ by jurisdiction, by channel, and by how the interface presented the agent to the buyer. Take that question to your own counsel with your actual message templates in hand. What is settled is enough to act on: automation is not a defence, no human needs to have read it, product promises create warranties, and deception is deception.

The Four Binding Sentences: Price, Terms, Timing, Claim

Here is the framework. Every sentence an AI sales agent produces falls into exactly one of five buckets. Four of them bind you.

One: Price. Any number attached to money the buyer would pay or save. List price, discount percentage, a waived fee, an extended trial, a credit, a price hold, a "we can probably get that approved." Price sentences are the highest-frequency binding class in practice because buyers ask about price constantly and models are eager to be helpful.

Two: Terms. Any statement of what the contract will say. Notice periods, auto-renewal, data residency, liability caps, termination for convenience, payment terms, the shape of an SLA. These arrive in email far more often than anyone expects, because a prospect asks "can we get out after year one?" and a helpful assistant answers the question rather than routing it.

Three: Timing. Any statement about when something will exist or arrive. Delivery dates, onboarding duration, implementation timelines, "that ships in Q4," stock availability, "we can have you live by the 15th." Timing sentences are the class most likely to be honestly believed and still wrong, because the agent is repeating a roadmap it read.

Four: Claim. Any assertion of fact about what the product does, how it performs, what it complies with, or who else uses it. Throughput numbers, uptime, certifications, integrations, named reference customers, security postures. This is the class that lands in UCC § 2-313 territory and in FTC territory simultaneously, and it is the one where a model's fluency is most dangerous, because a plausible capability sentence reads exactly like a true one.

Five: everything else. Scheduling, acknowledgement, summarisation, links to published material, restating what the buyer said, asking a qualifying question, logging a call. This bucket is enormous. It is most of the volume. It should move without a human.

The point of naming the classes is that a machine can check them cheaply, before send. You are not asking a classifier to assess risk, which it does badly. You are asking a shallow question with a crisp answer: does this sentence state a price, a term, a date, or a capability? That is closer to named-entity recognition than to judgment, and it fails safe if you tune it to over-flag.

Decision flow for the Honour Test: four questions applied to one AI-drafted sentence, ending in send, owner approval, or refuse

The Honour Test: A Diagnostic You Can Run in One Sitting

The Honour Test is one question, applied to one sentence: if the buyer screenshots this and holds us to it in six months, do we perform? If the answer is yes, or "we would have to," it is a commitment and it needs an owner. If the answer is "that is not a promise, it is a fact we can already evidence," it can go.

Run it over a real sample rather than in the abstract, because the abstract version always produces the answer "our agent doesn't do that." Here is the procedure, and it takes an afternoon.

Prerequisites. You need three things before you start, and if you cannot get them the exercise is not worth running. First, export access to the last 30 days of AI-assisted outbound and reply traffic — the actual sent bodies, not the templates. Second, one person who knows what the company will and will not honour: normally a deal desk lead, a sales operations manager, or in a smaller company the founder. Third, the current price list and the current standard contract, in the versions that were live during that 30-day window. Without the third item you will spend the afternoon arguing about whether a discount was in policy.

Step 1 — Pull 200 sent messages at random. Not the best ones. Random. If you have fewer than 200, take all of them and note the smaller sample in your write-up.

Step 2 — Split into sentences and drop the non-substantive ones. Greetings, sign-offs and calendar links are noise.

Step 3 — Classify each surviving sentence into one of the five buckets. Do this mechanically. A sentence that mentions a number and a currency is Price. A sentence that mentions a contractual mechanism is Terms. A sentence with a date or a duration attached to a company action is Timing. A sentence asserting a product fact is Claim. Everything else is bucket five.

Step 4 — For every sentence in buckets one to four, apply the Honour Test. Your deal desk lead answers yes or no. Record the count and, critically, record which ones the company would not have honoured. Those are your live exposures.

Step 5 — Count the approval load. Take the number of commitment-class sentences and divide by the number of messages. That ratio is the number your entire control design hangs on, and it is the number nobody measures.

We are not going to give you a benchmark ratio here. We have not run this on a client's outbound corpus, we have not found a published measurement of it, and we will not invent one to make the section feel more authoritative. What we will say is that this ratio, whatever yours turns out to be, is the only input that tells you whether commitment gating is cheap or expensive at your volume, and nobody we can find is measuring it. If it comes back high enough that gating looks unaffordable, read that as a finding about your agent's instructions before you read it as a finding about your control design: an agent that keeps reaching for discounts and delivery dates has been told to close rather than to converse.

Step 6 — Write the two lists. List A: the commitment classes this agent is allowed to produce, with a named owner per class. List B: the classes it may never produce, and what it says instead. That is your first Commitment Authority Card, and the rest of this article is about filling it in properly.

Nine Real Outbound Sentences, Classified and Ruled On

Abstract classes are easy to agree with and hard to apply. Here are nine sentences of the kind that actually appear in AI-assisted sales email, with a class and a verdict on each. We wrote these as representative templates rather than lifting them from any company's real correspondence.

#SentenceClassVerdict
1"Happy to send over the security overview — here is the link."Bucket fiveSend. Links to published material commit nothing new.
2"Our standard is $40 per seat per month, published on our pricing page."PriceSend only if the agent quoted from the live price list and the message cites it. Otherwise gate.
3"I can do 22% off if you sign for two years."PriceGate. Named owner, every time.
4"Yes, you can cancel with 30 days' notice."TermsGate unless a policy object says the standard contract has that clause and no rider applies.
5"We can have you fully onboarded before the end of the month."TimingGate. This is a resourcing promise dressed as a courtesy.
6"We're SOC 2 Type II certified."ClaimSend only if wired to a live attestation record with an expiry date. Otherwise gate.
7"We handle about 10,000 documents an hour at peak."ClaimGate. A performance number is an express warranty candidate under UCC 2-313.
8"Most teams your size see payback in under a quarter."ClaimGate, and probably refuse. This is an outcome prediction about the buyer's business.
9"Does Thursday at 2pm work for a 30-minute call?"Bucket fiveSend.

Sentences 2 and 6 show the mechanism that makes this design affordable. Neither needs a human. Both need a source of truth the agent must quote from and cite, rather than a memory it may recall from. Require the agent to retrieve the current price from the price list object and record the retrieval, and the sentence stops being a commitment the agent invented; it becomes a repetition of one the company already published. That converts a gate into a lookup, and lookups scale.

Sentence 8 is the one most teams will argue about. It reads like ordinary sales language, and under UCC § 2-313(2) a statement that is genuinely just the seller's opinion does not create a warranty. But "most teams your size see payback in under a quarter" is not obviously opinion; it is phrased as an observed distribution. If you do not have the data behind it, that sentence is an advertising-substantiation problem before it is a warranty problem. The FTC's Policy Statement Regarding Advertising Substantiation of 23 November 1984 requires that advertisers "have a reasonable basis for advertising claims before they are disseminated," because "objective claims for products or services represent explicitly or by implication that the advertiser has a reasonable basis supporting these claims." That is policy and case law rather than statutory text, which is why we cite the policy statement and not § 45. Our rule is simple: an AI agent may not make a statistical claim about outcomes the company has not published and cannot evidence on request.

What Is Actually Safe to Send Without a Human

The section nobody writes, and the one that makes the rest affordable. If the answer to "what can AI send unreviewed?" is "very little," nobody adopts the control, and you end up with the toggle flipped to send everything by the second week.

Safe to send with no human in the path, provided the agent's retrieval is scoped to published material:

  • Restating the buyer's own words. Call summaries, requirement recaps, confirmations of what was asked. These create no new obligation because the content originated with the other side. The derived record does still need governing, which we cover on conversation intelligence.
  • Scheduling and logistics. Meeting times, attendee lists, calendar links, reschedules.
  • Links to published artefacts. Pricing pages, documentation, case studies, security pages, terms of service. If it is already public, pointing at it adds nothing.
  • Qualifying questions. Budget, authority, need and timing questions ask rather than assert.
  • Internal CRM writes, if scoped. Notes, activity logging, enrichment, next-step fields. Permission scope matters far more than content here; a write to a closed-won amount field is not a note.
  • Acknowledgements and routing. "I have passed this to our security team, you will hear from Priya by Thursday" is a timing commitment about a colleague and should be gated; "I have passed this to our security team" is not.

That last pair is the whole discipline in miniature. Two sentences, one clause apart, one free and one gated. Any control design that cannot tell them apart is operating at the wrong resolution.

What We Found When We Checked Twelve AI Sales Products

We ran a small, reproducible check rather than asserting a market condition. On 31 July 2026 we fetched the public product pages of twelve AI sales and revenue tools and searched the rendered text of each for four things: any form of the word "approve," the phrase "human in the loop," any construction meaning review-before-send, and the word "guardrail."

ResultCount
Pages fetched12
Returned HTTP 20010 (Outreach's AI product URL 404'd; 6sense returned 403)
Contained any form of "approve"3 of 10
Contained "human in the loop"0 of 10
Contained "review before send" or "approve before"0 of 10
Contained "guardrail"3 of 10

Where "approve" did appear, it appeared once per page and in three different senses. HubSpot's AI page described approving automatic CRM updates and follow-up drafts after a meeting. Lindy's homepage stated that approvals are built in, in a security context, without describing their unit. Artisan's page was the most explicit, offering an "Approval mode" whose description is a binary: have the agent draft every reply and wait for your click, "or let her send." In fairness, the same list also offers escalation rules that hand off the moment a conversation crosses a line the customer sets, which is a real conditional control. It is a conversation-level control, though, not a sentence-level one, which is a different resolution from the thing this article argues for.

State the limits of that measurement plainly, because they are large. Marketing pages are not feature inventories; the absence of a word is not the absence of a capability, and several of these products very likely have approval mechanics documented behind a login we did not have. Two of the twelve URLs did not resolve at all. And a homepage scan is a shallow instrument by construction.

What survives those caveats is narrow but real: across the public-facing surface where these products explain what they do, the vocabulary of granular commitment control is essentially absent, and where approval is described, it is described as a per-campaign or per-reply switch. Nobody is talking about the unit. That is the gap this article is written into.

Why Approval Mode Is the Wrong Unit of Control

Most AI sales automation ships with a per-message approval toggle. It has two settings and both of them are wrong.

Set it to review everything and you have rebuilt the bottleneck the agent was bought to remove. A rep reviewing 400 drafts a week is doing a worse version of their old job, with the added indignity of reading someone else's prose. Review fatigue is not hypothetical; it is the predictable outcome of asking a human to approve a stream in which nearly every item is routine and nearly identical to the last one. After two weeks the reviewer is clicking approve without reading, which is worse than no gate at all because it manufactures an audit trail that says a human checked.

Set it to send everything and you have accepted, without deciding to, that your company will honour whatever a language model writes about price, terms, delivery and capability. That is not a risk appetite anyone signed off. It is a default.

The unit of control should be the commitment class, not the message. That gives you a third setting: send the routine traffic, stop the commitment sentences, and stop them at a named owner who has the authority to say yes. The reviewer is now looking only at sentences that genuinely need a decision, and each review is fast because the sentence arrives with its class and the relevant policy attached. How many that is per week is exactly the commitment-class ratio from the Honour Test, which is why the measurement comes before the design rather than after it.

This is the same design principle as approval-as-control in human-in-the-loop automation, applied at sentence resolution instead of workflow resolution. An approval step that fires on everything is not a control; it is a queue. An approval step that fires on the thing that actually creates obligation is a control.

Three implementation notes that decide whether this works:

  1. The gate must sit between draft and send, not between send and audit. A retrospective flag is a report, not a gate. If the message has left, the commitment exists and you are now negotiating your way out of it.
  2. The owner must be a role with real authority, resolved at gate time. "Sales manager" is not an owner. "The deal desk approver on rota for EMEA mid-market" is. If the gate routes to someone who cannot actually approve a 22% discount, they will approve it anyway and you have laundered the decision.
  3. A gate with no timeout policy will be bypassed. Decide in advance what happens when nobody approves within four business hours: the message goes without the commitment sentence, or it does not go. Both are defensible. Silence is not.

The Commitment Authority Card: One Agent, Fully Specified

Here is the worked artefact, filled in for a plausible agent so you can copy the shape rather than the content. This is the one document we think every AI sales agent should have before it sends a single external message, and it should fit on one page.

Agent: Mid-market inbound reply agent Owner (human, named role): Director of Sales Operations, EMEA Non-human identity: agent-inbound-emea-01, owner-scoped, expiry 2027-01-31 Channels: Email replies to inbound leads only. No outbound cold. No chat. No voice. Source-of-truth objects it must quote from: live price list PL-2026-Q3; standard MSA MSA-v9; published security page; published integrations list Retrieval rule: every commitment-class sentence must cite the object and version it came from, in the outgoing record

ClassAuthorityGateNamed approverFallback text if refused
Price — list priceAutonomous, must quote PL-2026-Q3Nonen/an/a
Price — any discountNoneHard gateDeal desk approver on rota"Let me get you an exact number from our deal desk today."
Terms — standard MSA clausesAutonomous, must quote MSA-v9 clause number and confirm no account rider modifies itGate if a rider existsLegal counsel"Your agreement has a rider on that clause; let me get the exact wording."
Terms — anything non-standardNoneHard gateLegal counsel"That one needs our legal team; I will come back to you by tomorrow."
Timing — published availabilityAutonomous, must quote published pageNonen/an/a
Timing — onboarding or delivery datesNoneHard gateImplementation lead"I do not want to guess at a date; our implementation lead will confirm."
Claim — published capabilityAutonomous, must quote docs URLNonen/an/a
Claim — performance numbers, certificationsNoneHard gateProduct marketing + security"Our security page has our current attestations; happy to walk through them."
Claim — outcome or ROI predictionProhibitedRefusen/a"I would rather show you what comparable teams measured than predict your number."
Bucket fiveAutonomousNonen/an/a

Disclosure line, always present: the agent identifies itself as an AI assistant in the first message of any thread. Record retained per message: prompt version, model and version, retrieved objects and versions, classifier output per sentence, gate decisions with approver identity and timestamp, final sent body. Review cadence: weekly for the first month, then monthly. Card version and effective date on the card itself.

Annotated layout of a Commitment Authority Card showing agent identity, source-of-truth objects, per-class authority rows and the retained record

The fallback text column is the part teams skip and the part that determines adoption. A gate that produces silence produces a rep who turns the gate off. A gate that produces a good sentence — one that keeps the conversation moving while the commitment is being approved — produces a rep who leaves it on. Write the fallback lines with the same care as the outreach copy.

Note what the card does with the non-human identity. The agent has an owner, a scope and an expiry, the same way a contractor's account would. We have written about treating non-human identities as first-class, and the sales case is the cleanest illustration of why it matters: when the Director of Sales Operations changes job, somebody has to inherit the answer to "who said this agent could offer 22%."

Disclosure: Telling the Buyer It Is a Machine

Separate from what the agent may promise is whether the buyer knows they are reading a machine. In the EU this stops being a matter of taste on 2 August 2026.

Article 50(1) of the EU AI Act requires that providers "ensure that AI systems intended to interact directly with natural persons are designed and developed in such a way that the natural persons concerned are informed that they are interacting with an AI system, unless this is obvious from the point of view of a natural person who is reasonably well-informed, observant and circumspect, taking into account the circumstances and the context of use" (Regulation (EU) 2024/1689). Article 113 of the same regulation sets the general date of application at 2 August 2026.

On scope: Article 50(1) is written around systems "intended to interact directly with natural persons," which plainly covers a live chat widget and an AI voice agent. Whether a one-way AI-drafted email that no human reviews is a "direct interaction" for these purposes is the sort of question that gets litigated rather than assumed, and the roles the regulation assigns to providers and deployers add a second layer to it. We are not going to give you an answer. Our read of the regulation's timeline sits in our EU AI Act guide for deployers; whether your outbound programme is in scope belongs with counsel who can see your actual flows.

On practice: the disclosure argument inside most sales organisations is not really a legal argument. It is a fear that saying "this is an AI assistant" kills reply rates. That may be true, it is measurable, and it is a decision the business is entitled to make with its eyes open. What it is not entitled to do is make the decision by never raising it. Put the disclosure question on the Commitment Authority Card and record who decided.

And what we are not telling you: whether a US state's unfair-practices law or a sector regulator currently requires AI disclosure in a B2B sales email. Several jurisdictions have bot-disclosure statutes with narrow scopes, the boundaries are contested, and we found no published enforcement on this fact pattern. Treat it as an open question rather than a settled duty.

The Record You Need When Someone Screenshots the Email

Assume the sentence went out and the buyer is holding you to it. What do you need to be able to produce, four months later, to make a decision about whether to honour it?

Five artefacts, and if you cannot produce all five, your default answer is going to be "honour it," because you will not be able to establish anything else.

ArtefactWhy you need itWhere it usually goes missing
The exact sent bodyEstablishes what was actually said, versus what the template saidProviders store the template, not the render
The retrieval recordShows which price list or contract version the sentence came fromNot captured at all in most stacks
The prompt and model versionEstablishes whether behaviour changed under youOverwritten on every prompt edit
The gate decision and approver identityDistinguishes "a human approved this" from "nobody saw it"Logged as a boolean with no identity
The refusalsShows the control was live and working, not just configuredAlmost never retained
Timeline from send to a customer screenshot four months later, with the five artefacts you must produce

The fifth row is the one that separates a real audit trail from a compliance screenshot. A log that records only what was sent tells you nothing about whether the gate was functioning; a log that also records what the agent tried to send and was stopped from sending is evidence of an operating control. That principle generalises well beyond sales, and it is the core of what we have argued about proving agent actions through audit trails.

One warning about retention. That record is high-volume and the sent bodies contain customer PII. Set the retention period when you decide to keep it, with the people who set your CRM retention policy, or you will have built a large, unclassified, indefinitely-retained store of customer correspondence in order to solve a governance problem. That is a trade, not a free win.

A Ninety-Day Sequence to Put Commitment Authority in Writing

Sequenced so each stage produces something usable even if the next never happens. That constraint matters more than speed.

Days 1 to 14 — Measure, do not design. Run the Honour Test over 200 real messages. Produce the commitment-class ratio and the list of sentences the company would not have honoured. Nothing else. If the exercise stops here you still have the single most useful number in the programme.

Days 15 to 30 — Write one card. Pick the agent with the highest external volume, not the most interesting one. Fill in the Commitment Authority Card, including fallback text. Get the deal desk and legal to argue about the rows now, on paper, rather than later, in a deal. Expect the argument about discount thresholds to take a full session on its own.

Days 31 to 50 — Wire the source-of-truth objects. Before any gate, make the agent retrieve and cite. This is the step that converts most Price, Terms and Claim sentences from commitments into repetitions, and it does more for your exposure than the gate does. It is also the step with real engineering cost, because "the live price list as an object the agent can query" does not exist in most companies.

Days 51 to 70 — Turn on classification in shadow mode. Classify every outgoing sentence, gate nothing, and compare the classifier's output to the human classification you did in days 1 to 14. You are looking for the false-negative rate on the four binding classes. Tune to over-flag; a false positive costs a reviewer ten seconds, a false negative costs a discount.

Days 71 to 90 — Enable gates, one class at a time. Price first, because it is the highest-frequency and the easiest to route. Then Claim, then Timing, then Terms. Watch the approval latency and the bypass rate. If either goes bad, the fault is almost always the owner definition or the missing timeout policy, not the classifier.

This is the same three-beat shape as the gateway rollout method we use internally — observe first, enforce second, optimize third — applied to a sales surface. Measure before you gate, gate before you tune.

Six Ways This Fails in the First Quarter

  1. The classifier is tuned for precision. Somebody optimises for fewer interruptions and raises the threshold. Recall on the binding classes drops, the gate stops firing on borderline discounts, and nobody notices because the dashboard shows a lower gate rate, which reads as success.
  2. The approver is a distribution list. Route a gate to a shared inbox and you have created diffusion of responsibility with a timestamp. Nobody approves, the timeout fires, the rep escalates, the gate gets an exception, and the exception becomes the path.
  3. The source-of-truth object goes stale. The price list object is wired in Q1; the actual price list moves to a spreadsheet in Q2. The agent now cites a version number for a wrong price, which is worse than citing nothing, because the citation makes it look verified.
  4. The fallback text is bad, so reps write around the agent. If the refused-commitment sentence is clumsy, reps stop using the agent for anything near a commitment and write those emails themselves, unlogged. Your gate is perfect and your coverage is zero.
  5. The card is never versioned. Discount authority changes in a meeting, the card is edited in place, and six months later you cannot say what the policy was on the day the sentence went out. Stamp the card version into the message record.
  6. Voice is forgotten. The email agent gets a card and the voice agent does not, even though a spoken discount is the same commitment with a worse record. Voice needs its own card and a transcript retention decision on day one.

When a Human Rep Still Wins

There are deals where none of this applies, because the right answer is not to put an agent in the path at all.

Anything with a bespoke commercial structure. If the deal shape is being invented in the conversation — usage-based pricing built for one customer, a co-development clause, an unusual liability arrangement — the commitment classes are not classes, they are one-offs. No card covers it.

Negotiations proper. Once both sides are trading concessions, the sentence-level frame is the wrong one; what is needed is a mandate with limits set before the agent speaks at all, which we cover on setting the negotiation mandate first.

Relationships where being written to by a machine is itself the damage. Some accounts, some sectors, some seniority levels. A legitimate judgment call. Measure it if you can; respect it if you cannot.

Anything where the buyer has already been burned. A prospect in a recovery conversation after a failed implementation is not a candidate for AI-drafted reassurance, however well the gate works.

Regulated selling with its own suitability rules. Financial products, insurance, healthcare-adjacent sales and anything with a statutory suitability or advice duty carry obligations that sit entirely outside this framework. The commitment classes do not cover suitability, and nothing here should be read as saying they do.

Where LeapForce Fits

The control this article describes has four parts: an identity for the agent with an owner and an expiry, a scoped set of things it may touch, a gate between draft and send that fires on the commitment classes, and a record that includes refusals. LeapForce is building that layer — one controlled layer for every AI tool, connector, model and agent, with access and identity treating non-human identities as first-class, connectors providing action-level scoping and human-in-the-loop gates, workflows carrying the approval gates and policy checks, and observability and audit recording what was refused rather than only what ran. LeapForce is in active development and each of those pages carries its own per-capability build status of live, in development or roadmap; ask us which is which for your use case rather than assuming all of it ships today.

What LeapForce does not do is classify your sales sentences for you or decide your discount authority; the commitment classes, the card, and the argument between your deal desk and your legal team are yours, and any vendor claiming otherwise is selling you a policy it has not read.

Where This Analysis Is Still Uncertain

Five things we do not know, stated plainly.

We have not run the Honour Test on a client's live outbound corpus. The procedure in this article is derived from the structure of the problem and from the sentence samples we constructed, not from a measured engagement. The commitment-class ratio is the number that would make this article much stronger and we do not have it. Treat step 5 as the exercise you run, not as a finding we are reporting.

The attribution question is genuinely open. Whether a given AI agent's statement is legally attributable to your company under E-SIGN, and how apparent authority applies to a system with no intent, is being worked out. We have cited what is settled and refused to resolve what is not. Nothing here is legal advice, and the jurisdiction-specific answers should come from your own counsel with your own templates in front of them.

Our twelve-product scan is a shallow instrument. Marketing pages, one fetch, four search terms, two URLs that did not resolve. It establishes what these products say on their public pages, not what they do. A proper version of that measurement would require trial accounts and documentation access.

The market data is contested and we could not reach the primary source. The Federal Reserve's April 2026 FEDS Note by Jeffrey S. Allen reports three surveys giving three very different answers: roughly 18% of US firms had adopted AI by year-end 2025 in the Census Bureau's Business Trends and Outlook Survey, about 41% of individuals reported work-related generative AI use in the Real-Time Population Survey, and the Atlanta Fed's Survey of Business Uncertainty put 78% of the labour force at firms that have adopted AI, with 54% at firms using LLMs. The note attributes the spread mainly to sampling and units of analysis. We wanted the Census Bureau's own May 2026 release for the sector breakdown, and census.gov blocked our fetches, so we are citing the Federal Reserve's summary of that survey rather than the primary release.

We could not verify the Chevrolet dealership chatbot incident against a primary source. It is widely cited as the canonical example of a bot agreeing to an absurd price, and the coverage we could reach was secondary reporting plus the contemporaneous discussion thread. We have therefore used it only as context for a practitioner quote and built no claim of ours on it; if you are citing it in an internal paper, get a primary artefact first.

One further disclosure about sourcing: Reddit was not reachable for this piece, so every practitioner voice quoted above comes from Hacker News. That skews technical for a topic whose primary audience is sales leadership, and the absence of a sales-operations voice is a real gap in the evidence rather than a stylistic choice. We also attempted a YouTube search for a conference talk on this exact question and the API's daily search quota was exhausted, so this article ships with no embedded video.

 FAQ

Frequently asked questions

Sometimes, and automation is not the reason you would escape it. Federal law at 15 U.S.C. § 7001(h) says a contract cannot be denied effect solely because an electronic agent was involved, provided the agent's action is legally attributable to the party to be bound, and UETA § 14 as enacted in most states says a contract may form even if no individual reviewed the terms. In Moffatt v. Air Canada, 2024 BCCRT 149, a Canadian tribunal rejected the argument that a chatbot is a separate legal entity and found negligent misrepresentation. Whether a specific statement by your specific agent is attributable to you is a contested question that depends on jurisdiction, channel and how the agent was presented. Ask your counsel; do not take an answer from a vendor page.

In the EU, from 2 August 2026, Article 50(1) of the AI Act requires that people be informed they are interacting with an AI system unless that is obvious to a reasonably well-informed observer, and Article 113 sets that date. Whether an AI-drafted email that no human reviewed counts as a direct interaction under that provision is not settled, and we are not resolving it here. Outside the EU, several jurisdictions have narrow bot-disclosure statutes and we found no published enforcement on a B2B sales email fact pattern. Practically: decide it deliberately, record who decided, and put the answer on the agent's authority card.

Under UCC § 2-313 an affirmation of fact or a promise about the goods that becomes part of the basis of the bargain creates an express warranty, and § 2-313(2) makes clear no formal words like "warrant" or "guarantee" are needed and no specific intention is required. Nothing in that text turns on who typed the sentence. The same subsection excludes statements merely of value or the seller's opinion, which is why "our platform is excellent" is different from "our platform processes 10,000 documents an hour." Treat any performance number, certification claim or integration assertion as warranty-shaped and gate it.

Separate quoting from deciding. An agent that retrieves the current published price from a versioned price-list object and cites the version in the outgoing record is repeating a commitment your company already made; that needs no human. An agent that produces any number not present in that object — a discount, a waived fee, an extended trial, a price hold — is creating a new commitment and must stop at a named approver. The engineering work is making the price list a queryable object with a version stamp, and in most companies that does not exist yet, which is why this looks like a governance project and is actually a data project.

Our list is short and we would defend it: outcome or ROI predictions about the buyer's business, any statistical claim about customer results the company has not published and cannot evidence on request, any non-standard contract term, and any statement about a competitor's product. The first two are advertising-substantiation exposure before they are anything else, under the FTC's long-standing reasonable-basis requirement. The third belongs to legal by definition. The fourth generates disputes with no upside. Everything else can be either autonomous with a retrieval requirement or gated to a named owner.

We cannot give you a dollar figure, and any figure quoted without seeing your stack is invented. The cost has three parts whose relative sizes are stable even where the absolute numbers are not. The classifier is cheapest: sorting a sentence into five buckets is a small-model job. The source-of-truth objects are the expensive part, because making price lists, contract versions and attestation records machine-queryable is real engineering. The approval load is the ongoing part, and its size is exactly the commitment-class ratio the Honour Test measures.

Wrong question if you are asking it about the gating. The gate does not pay back; it prevents a loss you cannot forecast, in the same way a contract review does not pay back. The AI sales agent itself may well pay back quickly, and that number belongs to your own before-and-after measurement of rep hours and reply rates, not to a vendor's benchmark. What we would put a timeline on is the artefact: two weeks to a measured commitment-class ratio, four weeks to a written authority card, ninety days to gates live on all four classes.

Buy the drafting, own the gate. The drafting, sequencing, enrichment and deliverability work is a commodity that several vendors do better than you will, and there is no strategic value in owning it. The commitment classes, the authority card and the record are policy about your own company's obligations, and outsourcing those to whichever vendor happens to be in the outbound seat means re-deciding them at every renewal. In our public scan of twelve product pages, none described approval at the resolution this needs, so at present the gate is something you will be assembling regardless.

The named owner on the agent's authority card, which is the reason the card exists. If no card exists, the honest answer is that ownership will be assigned after the fact by whoever is in the room, which is how the same incident produces a fired rep at one company and a policy change at another. Name the owner per commitment class, not per agent, because the person who should approve a discount is rarely the person who should approve a security claim. And record the approver identity in the message record, not as a boolean, or you will be unable to reconstruct who said yes.

Yes, and more urgently, because a spoken commitment leaves a worse record and travels faster. The four classes are identical on a call: a quoted discount, a contract term described aloud, a delivery date, a performance claim. What changes is that a pre-send gate is not available in the same form, so the control has to move to what the agent is allowed to say at all, plus a transcript you can actually search. Recording-consent law varies by jurisdiction and is genuinely restrictive in some of them; get that specific question answered by counsel before you deploy voice, not after.

Ready to Govern Your AI?

Talk to LeapForce — one controlled layer for every AI tool, connector, model, and agent.

Thirty minutes · No pitch deck

Ready to turn AI experiments into measurable ROI?

Bring one outcome you'd like AI to move. We'll help you scope a pilot you can actually measure — and tell you honestly if it's not worth doing yet.

Comments