Conversation Intelligence: Govern the Scores, Not the Call

Conversation intelligence is software that records business calls, transcribes them, and then writes new records about the people who were on the call: talk-tim

Conversation intelligence is software that records business calls, transcribes them, and then writes new records about the people who were on the call: talk-time counts, topic labels, sentiment readings, deal-health scores. The recording is the part buyers negotiate over. The new records are the part that outlives the deal.

Our position is that the industry has the risk model backwards. Almost every conversation intelligence evaluation we have seen treats the audio file and the transcript as the sensitive asset and the analytics as the product benefit. It is the other way round. The transcript is a copy of something that already happened and can be checked against reality. The analytics are a manufactured record about a named employee and a named customer, produced by a machine, held to no accuracy standard, distributed into systems your retention policy has never heard of. In the European Union, one whole class of it has been outright prohibited in the workplace since February 2025.

On Hacker News, a commenter writing as tbragin put the discomfort precisely while discussing AI-assisted sales: "Previously it was just for note-taking, but now it can be much more powerful" (Hacker News, 29 September 2023). They were not objecting to being recorded. They were objecting to what the recording gets turned into, and they named sentiment analysis from voice as the line. That is the right instinct, and it is the one almost no buying process is built to act on.

The short answer: Buy conversation intelligence for what it can count, use what it can quote, and treat everything it infers, from sentiment and intent to coachability and deal health, as an unregulated personal-data pipeline that you must be able to name, date, justify and delete before it reaches a manager's dashboard.

Last updated: July 30, 2026.

We have not run a controlled deployment of a commercial conversation intelligence platform ourselves, so nothing below is presented as a measured result from our own rollout; every number is sourced and linked, and the worked example is labelled as illustrative arithmetic rather than a customer case.

Three tiers of derived record produced from one recorded call: observed counts, attributed extractions, inferred states

The three tiers of derived record. Retention rules are written about the call; the records that travel are the ones on the right.

What conversation intelligence actually is, and what it manufactures

Conversation intelligence is a category of software that captures spoken business conversations — sales calls, support calls, meetings — converts them to text, and then runs analysis over that text and audio to produce structured records: who spoke for how long, which topics and competitors came up, whether the customer sounded satisfied, whether the deal is progressing. It sits between a call recorder, which only preserves, and a CRM, which only stores what a human typed. The category overlaps heavily with call analytics in contact centres and with AI sales coaching on the revenue side; the underlying pipeline is the same in all three, and so is the problem this article is about.

The pipeline has six stages, and buyers routinely evaluate only the first two.

StageWhat happensWhat it producesWho checks it
CaptureThe platform joins the call or ingests the telephony streamAudio fileAudio quality is obvious
TranscribeSpeech recognition converts audio to textTranscriptReps notice bad transcripts immediately
DiariseThe system separates and labels speakersSpeaker-attributed transcriptRarely audited
CountDeterministic metrics over the diarised transcriptTalk ratio, monologue length, question countAlmost never audited
ExtractNamed entities, topics, objections, commitmentsTopic tags, competitor mentions, next stepsAlmost never audited
InferModels assert unobservable statesSentiment, intent, deal health, coaching scoresEffectively never audited

Everything from "Diarise" downward is derived data. It did not exist before your platform created it. It is about identified people. And it is generated at a volume no human reviews: a 60-rep team on four calls a day generates roughly 5,000 calls a month, each carrying dozens of derived fields.

What conversation intelligence is not. Three confusions cost real money in evaluations:

  • It is not conversational AI. Conversational AI holds the conversation; conversation intelligence watches one that humans are already having. We covered the custody problem on the first of those in our earlier analysis of who keeps the transcript, and the distinction matters here because the two categories have opposite failure modes: a conversational AI says something wrong to a customer, while conversation intelligence says something wrong about an employee.
  • It is not a call recorder with a search box. A recorder's output is a copy. This category's output is a judgement.
  • It is not a compliance control, even when it is sold beside one. Detecting that a disclosure was missed after the call ended is not the same as preventing the call from proceeding without it.

The three tiers of derived record

Sort every field a conversation intelligence platform produces into one of three tiers, by the question "what would it take to prove this wrong?" The tiers are not a maturity model. They are three different kinds of claim with three different duties attached, and mixing them on one dashboard is where governance breaks.

TierNameExample fieldsHow you falsify itGovernance duty
1ObservedTalk ratio, longest monologue, question count, interruptions, call durationReplay the call and recountPublish the definition; the metric is only as good as the diarisation under it
2AttributedCompetitor named, pricing objection raised, next step agreed, discount promisedFind the line in the transcriptShow the source line beside every tag; a tag without a quote is not evidence
3InferredSentiment, emotion, buying intent, deal health, rep coachability, churn riskYou cannot — the claim is about a mindTreat as a new personal record with a subject, a purpose, a retention clock and a contest path

Tier 1 is measurement. You can argue about whether a 62/38 talk ratio means anything for win rates, but you cannot argue about the number once the diarisation is right.

Tier 2 is extraction. It inherits every error from the layers beneath it. If the transcript renders "we're already talking to Acme" as "we're already walking to Acme", the competitor-mention tag silently disappears — but it is still checkable in seconds, because a correct tier-2 field can always point at a line.

Tier 3 is the one this article exists for. There is no line to point at. "Customer sentiment: negative" is not present in the conversation; it is asserted about the conversation. When a rep disputes it, there is nothing to look up. When a regulator asks how it was produced, there is a model. When the customer asks what personal data you hold about them, that score is personal data about them and you have to disclose it.

The single most useful thing an evaluation can do is ask the vendor to hand over a list of every field the platform writes, and sort that list into these three tiers before any pricing conversation starts. Across the vendor documentation we reviewed for this piece, we did not find a published field inventory in a form that makes this sorting easy. Expect to build it yourself from product documentation and a trial tenant.

In the European Union, the part of conversation intelligence that infers the emotions of workers is not a quality problem to be improved. It is prohibited. Article 5(1)(f) of the EU AI Act bans "the placing on the market, the putting into service for this specific purpose, or the use of AI systems to infer emotions of a natural person in the areas of workplace and education institutions, except where the use of the AI system is intended to be put in place or into the market for medical or safety reasons." That prohibition is not on a future timetable: Article 113 states that Chapters I and II, which contain the prohibited practices, apply from 2 February 2025.

Three details in the definitions do most of the damage to the "our tool is different" defence.

First, the scope is wider than faces. Article 3(39) defines an emotion recognition system as one "for the purpose of identifying or inferring emotions or intentions of natural persons on the basis of their biometric data," and Article 3(34) defines biometric data as personal data from technical processing of "physical, physiological or behavioural characteristics." Voice is a behavioural characteristic. A tool that reads emotional tone from vocal features is squarely inside the definition, and the inclusion of intentions reaches further than most vendors' marketing has noticed.

Second, the prohibition attaches to the area, not the subject. It applies in the workplace. A sales call has a worker on one side of it. The Future of Privacy Forum's analysis of the prohibition notes that voice analysis can constitute biometric data and points at the Hungarian enforcement action as the precedent for exactly this pattern (FPF, red lines under the EU AI Act). The FPF analysis also draws boundaries that matter in the other direction: physical states such as fatigue or pain are excluded, as is emotion inference from written text. So a platform that scores sentiment from the transcript rather than the voice signal sits in a materially different position from one reading vocal affect, and is worth asking about explicitly.

Third, the penalty tier is the top one. Article 99(3) sets fines for breaching the Article 5 prohibitions at up to EUR 35,000,000 or 7% of total worldwide annual turnover, whichever is higher. That is the highest band in the regulation.

This is not hypothetical enforcement waiting for the AI Act to warm up. A closely analogous case was decided years earlier under the GDPR, on a different legal basis and before the AI Act existed — which is why it is evidence about regulator appetite rather than about how Article 5 will be applied. In 2022 the Hungarian data protection authority fined a bank HUF 250 million (about EUR 665,000) over AI voice analysis applied to customer service call recordings, and — the part worth reading twice — found the solution "inefficient in predicting the customers' emotions accurately" (DLA Piper Advocatus). The regulator did not merely object to the processing. It examined the accuracy of the inference and found it wanting, then ordered the emotion analysis stopped.

The United States problem is different and arrives sooner

In the US there is no equivalent prohibition, and the pressure comes from wiretapping statutes instead, with the striking feature that the exposure attaches to the vendor's capability, not to what the vendor actually did.

Under a line of California Invasion of Privacy Act cases, providers of call-analysis software can be exposed as third-party listeners for "maintaining the mere capability to use the contents of the collected communications for their own purposes, such as training their AI, regardless of actual use," a reading that traces to Javier v. Assurance IQ, LLC and was applied in Ambriz v. Google, LLC (N.D. Cal., 10 February 2025), where the court rejected Google's argument that its Contact Center AI was merely a tool used by its business customers (Debevoise Data Blog, CIPA litigation trends).

Read that as a procurement instruction. The contractual promise "we do not use your data to train our models" is the promise buyers ask for. The reading that has been surviving motions to dismiss asks a harder question: could the vendor, architecturally? A tenant-isolated deployment where the vendor has no read path to your audio answers a question a data processing agreement only makes a promise about.

The parallel exposure runs through biometric privacy law. Where a platform performs speaker recognition rather than merely speaker separation, voice moves from "recording" to "biometric identifier" in states with biometric statutes, and consent obligations attach per person on the call rather than per account holder. This is the difference between diarisation ("speaker A and speaker B") and voice identification ("this is the same speaker as in call 4,112"), and it is a question most evaluations never think to ask because the two features look identical in a demo.

JurisdictionThe binding questionThe wrong assumption
EUDoes anything in the pipeline infer emotion or intention from the voice signal, in a workplace?"It's just analytics" — the AI Act names inference, not analytics
CaliforniaCould the vendor use call contents for its own purposes, architecturally?"Our DPA forbids it" — capability has been the tested question
Biometric-statute statesDoes the platform identify speakers, or only separate them?"Diarisation isn't biometrics" — sometimes true, sometimes not
EverywhereWho is the subject of each derived field, and were they told?"The account holder consented" — the rep and the customer are separate subjects

The accuracy question, asked in the right order

Most accuracy discussions about conversation intelligence stop at transcription word error rate, which is the one number that matters least, because transcription errors are visible and every other error compounds on top of them.

Start with the floor. In a peer-reviewed study of five commercial speech recognition systems from Amazon, Apple, Google, IBM and Microsoft, Koenecke and colleagues measured "an aggregate WER of 0.35 (SE: 0.004) for black speakers versus 0.19 (SE: 0.003) for white speakers." Error rates ran roughly twice as high, in every system tested (Koenecke et al., PNAS, 2020, via PubMed Central). Speech recognition has improved since, and that study used curated interview audio rather than telephony. But the finding that matters is not the absolute figure; it is that the gap was present in all five systems, which means the derived record built on top of a transcript is not uniformly accurate across the people it describes.

Then note that real call audio is worse than study audio. Voicegain, an ASR vendor publishing its own 2025 benchmark on forty call-centre files, describes the conditions plainly: telephony is usually encoded at 8 kHz narrowband, contact-centre recordings carry "significant background noise and over-talk," and agents and customers speak with a wide range of accents (Voicegain 2025 call-centre benchmark). That is a vendor source and should be weighed as one, but the description of the conditions is not contested by anybody.

Now stack the layers.

LayerTypical framingThe error it addsVisible to the user?
Transcription"95% accurate"Words wrong, and unevenly across speakersYes — reps read transcripts
DiarisationRarely quoted at allWords attributed to the wrong speakerOnly if you look
CountingPresented as exactInherits every diarisation errorNo — a talk ratio has no error bar
Extraction"AI-detected topics"Missed and hallucinated tagsOnly if a quote is shown
Inference"Sentiment: negative"Unbounded — the construct itself is contestedNo

The last row deserves its own footing, because the problem is not model quality. It is that the thing being measured may not be measurable the way the product assumes. The most-cited review of the science, Barrett and colleagues in Psychological Science in the Public Interest, concluded after surveying the evidence that how people communicate anger, disgust, fear, happiness, sadness and surprise "varies substantially across cultures, situations, and even across people within a single situation," and that "similar configurations of facial movements variably express instances of more than one emotion category" (Barrett et al., 2019, via Europe PMC). That review addresses facial movements rather than voice, and it does not say emotional expression carries no signal. It says the mapping from expression to emotional state is not the reliable, universal one that commercial applications assume. A sentiment score built on that assumption does not become correct with more training data; it inherits a contested construct.

The practical consequence. A tier-1 metric that is 3% wrong is a rounding error. A tier-3 score that is 30% wrong, presented without an error bar next to a named employee's photograph, is a performance management document. The presentation layer erases exactly the distinction that matters most, and it does so by design: dashboards look better when every number has the same visual weight.

Where the derived record goes after it leaves the platform

Here is the failure most retention policies contain. Retention rules get written about the recording, and the derived records escape through a door the policy never mentions: integration.

Take a concrete, documented example. Gong's published retention policy states that its "standard data retention period for existing customers is the lesser of three years and the period you are a customer," and that if you stop being a customer your data is "irreversibly deleted in 30 days" (Gong data retention policy). Its deletion documentation is equally clear about scope: deleting a call removes "all call-related data, such as transcripts, media," along with notes, comments and call stats, and "deleted calls cannot be retrieved" (Gong, delete a recorded call). That is better behaviour than many platforms, and it should be said plainly: the derived stats are deleted with the call, inside the platform.

Inside the platform. The same category of product is bought precisely because it writes outward: into CRM fields, into chat channels, into the analytics warehouse, into dashboards and exports and scheduled reports. Every one of those writes creates a copy the platform's delete button does not reach.

DestinationWhat lands thereWhose retention rules applyReached by the platform's delete?
CRM opportunity recordDeal health, next steps, competitor tagsThe CRM'sNo
Chat channel postCall summary, sentiment call-outThe chat tool's, often foreverNo
Analytics warehouseEvery derived field, per call, per repThe data team'sNo
Coaching dashboard / scorecardAggregated rep scores over timeUsually nobody'sNo
Manager's exported spreadsheetWhatever was on screen that dayNone whatsoeverNo
Vendor model trainingAggregate patternsThe vendor'sDepends entirely on contract and architecture

Notice the asymmetry. The audio file has a retention policy, an owner, an access log and a delete button. The sentence "Priya's calls trend negative in Q3" has none of those, and it is the one that will be in a spreadsheet during a performance review eighteen months from now.

This is where conversation intelligence differs from the transcript-custody problem we analysed for conversational AI platforms. There, the question is who holds a copy of something that was said. Here, the question is who holds a copy of something the machine decided, and it is harder, because a derived record has no natural home. Nobody feels like its owner. A recording obviously belongs to the recording system. A number belongs to whoever is looking at it.

The governance move is unglamorous and it works: make the derived record a first-class object with an owner, the same way we treat non-human identity as first class: a thing with a name, a responsible human, a scope, and an expiry date. Data that nobody owns is data nobody deletes.

The Derivation Ledger: one page, seven columns

Give the derived record a name and it becomes governable. We call the artifact the Derivation Ledger: one row per field the platform produces, seven columns, filled in once during evaluation and reviewed when the vendor ships new fields.

The point is not the document. The point is that four of the seven columns are usually impossible to fill in, and discovering which four, before signing, is the entire exercise.

ColumnQuestion it answersWhy it bites
FieldWhat is this thing called on the screen?Vendors rename fields between releases; the dashboard label is what a manager will quote
TierObserved, attributed, or inferred?Determines every other column
Source signalTranscript text, audio features, CRM data, or a mix?Audio-feature emotion inference is the prohibited class in the EU; transcript-derived sentiment is not the same object
SubjectWhich named person is this a record about?Most fields are about two people at once, and only one of them is your employee
ConsumerWhich human decision does this feed?A field nobody acts on is a liability with no benefit — turn it off
ClockHow long is it kept, in each system it reaches?The answer is per-destination, not per-platform
Contest pathHow does the subject see it and challenge it?If there is no answer, you have built an unappealable record

A filled row looks like this:

FieldCustomer Sentiment (call-level)
Tier3 — inferred
Source signalTranscript text plus prosodic audio features (vendor to confirm in writing)
SubjectThe customer, and by implication the rep whose call it scores
ConsumerWeekly pipeline review; feeds the at-risk deal list
Clock3 years in the platform; indefinite in the CRM field it writes to; indefinite in the warehouse
Contest pathNone currently — no rep-facing view, no correction workflow

Two of those seven rows are red on the day you write them, and both are fixable before go-live: pin the clock by turning the CRM write off or adding a purge job, and create a contest path by giving reps read access to their own derived fields. The third, "vendor to confirm in writing" on source signal, is a question you can only ask before you sign, which is the argument for building the ledger during evaluation rather than after the rollout.

A rule that travels: no tier-3 field reaches a manager's screen until its Consumer and Contest path columns are filled. Fields with no decision attached get switched off, which is also the cheapest performance improvement available in this category. Every field you disable is one fewer number a manager can misread.

A worked example: one 38-minute call, traced end to end

The following is illustrative arithmetic built from the sourced figures above, not a customer engagement. It exists to show how the layers compound, and every input is labelled.

A 38-minute discovery call. Two speakers. Roughly 5,500 words spoken, which is ordinary for the length at conversational pace.

Layer 1 — transcription. Assume the favourable end of the range for clean audio and take a 5% word error rate. That is 275 wrong words in the transcript. Now apply the Koenecke finding: the error rate is not uniform across speakers, and in that study it ran roughly twice as high for one group as another. On this call, one speaker's transcript may be materially worse than the other's, and nothing in the interface says so.

Layer 2 — diarisation. Suppose 2% of utterances are attributed to the wrong speaker, which is unremarkable when two people talk over each other. On a call with maybe 180 speaker turns, that is three or four turns landing on the wrong person.

Layer 3 — counting. The reported talk ratio is 62/38. Those three or four misattributed turns move it by a couple of points in whichever direction the errors happened to fall. The dashboard shows "62%" with no interval. The coaching rule of thumb says "talk less than 55%." The rep is now on the wrong side of a threshold by an amount smaller than the measurement error.

Layer 4 — extraction. The customer said "we're already talking to Acme." If that phrase falls inside the 275 wrong words, the competitor-mention tag never fires. The forecast review that filters on "deals with a competitor mentioned" will not surface this deal. Nobody will ever know the tag was missing, because absence leaves no trace.

Layer 5 — inference. The customer was terse for the last ten minutes because their next meeting started early. The sentiment model reads terse as negative. "Customer sentiment: negative" is written to the opportunity record.

What happens next is the actual harm, and it is organisational rather than technical. The deal-health field flips to at-risk. The weekly pipeline review, which reads the field rather than the call, deprioritises it. The rep's quarterly summary shows two negative-sentiment calls out of nine. In the performance conversation, the manager is not making a judgement about the rep. They are reading a number that a machine produced from a terse afternoon, and the rep has no way to point at the line that produced it, because there is no line.

Contrast the tier-1 path on the same call. "Longest monologue: 4 minutes 10 seconds" is checkable in thirty seconds. The rep can listen to it, agree it was too long, and change something on Monday. That is the whole promise of the category, working exactly as advertised, and it does not require a single inference.

What conversation intelligence is genuinely good at

An article that only lists risks is not useful to somebody who has to make a decision this quarter. The category earns its price on four jobs, and all four sit in tiers 1 and 2.

Making calls searchable. Before this software, institutional memory of what customers said lived in the heads of people who left. Full-text search across every call your company has had is a genuine capability change, not an incremental one, and it does not require a single inferred field.

Turning coaching from opinion into evidence. A manager saying "you talk too much on discovery calls" is an assertion. A manager and a rep listening to a four-minute monologue together is a shared observation. The mechanism that creates value here is the timestamp, not the score.

Detecting the absence of things. Some of the most valuable checks are negative and deterministic: the required disclosure was not read, the next step was not agreed, the decision-maker was never on a call in a six-month cycle. These are extraction questions with checkable answers, and they catch process failures that no amount of sentiment scoring would.

Compressing onboarding. A new rep can hear twenty real calls in their first week instead of shadowing four. The value comes from the library, not the analytics.

Notice what is common to all four: the software's job is to make a human look at the right thirty seconds. Every one degrades when the product tries to replace the looking with a number.

This is also why the honest version of AI sales coaching is not a score. It is a queue of timestamps a manager and a rep listen to together, ranked by anything you like, including a tier-3 signal, used as a sorting hint that nobody records. A sentiment model is a defensible way to decide which nine of five thousand calls get human attention this week. It is not a defensible way to decide anything about the person on them.

Buying questions that actually separate vendors

Feature lists in conversation intelligence software converged years ago. Every platform transcribes, diarises, tags topics, scores sentiment and writes to a CRM. The questions below separate vendors because the answers vary, and because several of them cannot be answered from a demo.

#QuestionWhat a good answer looks likeWhat a bad answer sounds like
1Give us the complete inventory of fields you write, and tier each one.A document, before the trial ends"It depends on your configuration"
2For each inferred field: transcript text, audio features, or both?Field-by-field, in writing"Our model considers many signals"
3Can we disable individual inferred fields, per team and per region?Yes, with a settings screen we can see"You can hide them in the UI"
4Do you perform speaker identification, or only speaker separation?A clear architectural answerConfusion between the two terms
5Architecturally, could your staff or your models read our call contents?Tenant isolation described concretely"Our DPA prohibits it"
6When we delete a call, what happens to fields already written to our CRM?An honest "nothing — that's yours to purge""Everything is deleted"
7Can a rep see every derived field about themselves, and dispute one?A rep-facing view exists"Managers can share feedback"
8What is your EU position on Article 5(1)(f) for workplace emotion inference?A specific, dated, documented position"We're fully compliant"
9Show us accuracy figures for diarisation and extraction, not just transcription.Numbers with methodologyOnly a transcription percentage
10What is the export format, and can we get everything out including derived fields?A tested export"Contact support"

Question 6 is the one we would put first if only one could be asked. The honest answer is "nothing, because that field belongs to your CRM now." A vendor who says that has told you something true about how the category works. A vendor who claims full deletion across your CRM has either built something unusual or has not understood the question.

What it costs, above and below the sticker price

Published per-seat pricing exists at the self-serve end of this market and disappears at the enterprise end. Prices below were fetched on 30 July 2026 from vendor pricing pages and will drift.

ProductCategoryPublished price
FirefliesMeeting recording, transcription and summary, with integrationsFree tier; Pro $10/seat/month billed annually ($18 monthly); Business $19 annually ($29 monthly); Enterprise $39 annually (published)
OtterMeeting recording, transcription and AI workflowsPro $16.99/user/month ($8.33 billed annually); Business $30/user/month ($19.99 billed annually) (published)
GongFull revenue-focused conversation intelligence platformNot published. The pricing page states licences are priced per user, with a platform fee based on the number of users supported, and directs buyers to request a quote (quote-based)

The first two are transcription-and-summary tools that many teams use as an entry point; the third is a full platform of the kind this article is about. They are not equivalent purchases, and the price gap between the rows is mostly the derived-record machinery, which is another way of saying you can buy the searchable library cheaply and the inference expensively.

The visible price is the smaller half of the number for anyone deploying this across a regulated team. The costs below are not on any pricing page and are the ones that decide whether the rollout survives its second year.

Hidden costWhere it landsWhy it is unavoidable
Field inventory and tieringAn analyst-scale exercise, once, plus a review each time the vendor ships new fieldsNobody publishes it for you
Consent and disclosure changesLegal review, call-opening script changes, regional variantsConsent rules differ by jurisdiction and by subject
CRM write-back purge jobsEngineering, ongoingThe platform's delete does not reach your CRM
Subject access requestsSupport and legal, per requestDerived scores are personal data about the person they describe
A contest path for repsProduct configuration plus a documented processWithout it every score is unappealable
Manager training on tiersRecurring, because managers rotateThe tiering only works if the people reading the dashboard know it exists

We cannot put a defensible dollar figure on that column without inventing one, and we will not. What we can say is which line items are recurring rather than one-off: the field review, the subject access requests, the purge jobs and the manager training all repeat, while the inventory and the consent redesign are largely front-loaded. That shape is enough to build a business case honestly; the amounts have to come from your own legal and engineering rates.

One implementation detail carries most of the CRM purge cost, and getting it right at go-live is far cheaper than retrofitting it. Write every derived field into a dedicated, clearly named field group on the CRM object, and stamp each one with the source call ID. Purging then becomes a scheduled job over one field group keyed to calls that have aged out, rather than an archaeology project across fields that three teams now depend on. Platforms that write into your existing standard fields by default make this materially harder, which is worth knowing during evaluation rather than during the first deletion request.

There is a genuinely cheaper option, and it is worth costing before the full deployment: buy the recording, transcription and search, and turn off every tier-3 field. You keep the searchable library, the coaching timestamps and the negative-check extraction, and you delete an entire category of legal exposure and appeal process.

Be clear about the trade, because it is real. You lose automatic at-risk deal flagging, which some forecasting processes now depend on. You lose the churn-risk queue that service teams use to triage. And you may lose the ability to buy the product at all in that shape. Several platforms bundle inferred scoring into the tier that also contains the features you want, and per-field disable switches are not universal. Confirm that switch exists before you plan around it. Where it does, this is most of the value at a fraction of the governance cost, and it is a defensible answer to "we already bought the licences" — you do not have to switch vendors to switch postures.

Where LeapForce fits, and where it does not

LeapForce does not sell conversation intelligence, call recording, or sales coaching software, and nothing on this page should be read as a substitute for evaluating the platforms above on their merits. What we build is the layer underneath: one controlled place where AI tools, connectors, models and agents are given identity, scope, policy and an audit trail, so that when a system starts generating records about your employees and customers, somebody owns it, its reach is bounded, and what it did is provable. The relevant pieces here are observability and audit, which is built to record what was refused and not only what ran, and the connector registry, where the write paths a tool is allowed to use are scoped once by IT rather than negotiated per team. Our gateway rollout guide follows a sequence we would apply to this category unchanged: observe first, enforce second, optimize third — point traffic at the gateway in observe mode and find out what is actually being generated and where it goes, before writing a single policy about it.

Honest limits and open questions

Several things in this analysis are genuinely uncertain, and a reader making a decision deserves to know which.

We have not tested these platforms. This piece is built from vendor documentation, published law, peer-reviewed research and litigation coverage. It contains no measured deployment of our own, and where accuracy figures appear they come from studies whose conditions differ from your call audio.

This is general information about a regulated area, not legal advice. Every jurisdictional line described here needs to be checked by counsel against your actual deployment, your actual call flows and the states and countries your customers sit in.

Two of the load-bearing research citations are several years old. The Koenecke speech recognition study dates from 2020 and the Barrett emotion review from 2019, and models have improved substantially since both. We cite them anyway because neither finding is about a model generation: one is about error being unevenly distributed across speakers, the other about whether the underlying construct maps cleanly to observable signals. Both would need to be refuted rather than merely outdated, and we found no published work doing so.

The AI Act line between prohibited emotion inference and permitted analysis is not settled. The FPF analysis makes clear that mood is distinguished from emotion, that physical states are excluded, and that emotion inference from written text sits outside the definition. But the boundaries require case-by-case judgement, and a platform whose sentiment scoring runs on transcript text is in a different position from one reading vocal affect. If your vendor cannot tell you which it does, that is the finding.

The CIPA capability line is live litigation, not settled law. The Debevoise summary describes courts splitting between an "extension" approach and a "capability" approach to whether an analytics provider is a third party. Ambriz was a motion-to-dismiss ruling, not a final judgement, and the doctrine may narrow.

We could not verify the efficiency claims the category markets on. Time-saved and win-rate figures are widely quoted in vendor material with no published methodology, and we found none that reconciled to a documented study, so none appear here. If a vendor shows you one, ask for the sample, the period and the counterfactual before it enters a business case.

Reddit and several practitioner communities were unreachable for this research, so the field voices here come from Hacker News, which skews technical for a topic whose primary buyers are sales and customer-success leaders. Treat the problem framing as representative of engineering-adjacent concern rather than of the whole buyer population.

We found no credible video on the specific question of governing derived conversation records, so none is embedded. The YouTube Data API quota was exhausted during this research and the search could not be completed to our normal standard, which is a limitation of this run rather than a claim that nothing exists.

Where this analysis could be wrong. If sentiment inference on business calls turns out to be substantially more reliable than the emotion-science literature suggests — a claim that would need published, independent, per-vendor accuracy figures, which do not currently exist — then the tier-3 caution here is overpriced, and the correct posture is closer to what vendors already sell. We would change our position on published evidence. We have not seen any.

 FAQ

Frequently asked questions

No, and confusing them causes real evaluation errors. Conversational AI holds a conversation with a person — it generates the words. Conversation intelligence watches a conversation two humans are already having and produces records about it. Their failure modes are opposites: conversational AI can say something wrong to a customer, while conversation intelligence can say something wrong about an employee. Different buyer, different risk register, different vendors.

Recording and transcribing calls, and analysing what was said, remains lawful with a valid GDPR basis and proper notice. Inferring emotions is different. Article 5(1)(f) of the EU AI Act prohibits using AI systems to infer emotions of a natural person in the workplace, except for medical or safety reasons, and per Article 113 that prohibition has applied since 2 February 2025. Because a sales or support call has a worker on one side of it, a platform that scores emotional tone from the voice signal needs a documented position. This is general information, not legal advice. Get counsel on your specific deployment.

Consent rules vary by jurisdiction and, critically, by subject. The customer, the employee and any third party who joins the call are separate data subjects with separate rights. In the EU, the Hungarian regulator held that emotion-based voice analysis required the freely given, informed consent of the data subjects. In the US, recording consent depends on state law, and a separate exposure runs to the software vendor itself under California's wiretapping statute. Two-party-consent states and multi-jurisdiction teams need per-region call-opening scripts, not one global disclosure.

No platform we found publishes independently verified accuracy figures for sentiment on business calls, and the deeper problem is that the construct itself is contested. The most-cited scientific review, Barrett and colleagues in Psychological Science in the Public Interest, found that how people express the common emotion categories "varies substantially across cultures, situations, and even across people within a single situation." Treat any sentiment number as a weak signal for prioritising which calls a human should listen to, and never as evidence in a performance decision.

Self-serve transcription and meeting tools publish per-seat prices: as of 30 July 2026, Fireflies lists $10 to $39 per seat per month billed annually and Otter lists $8.33 to $30 per user per month depending on tier and billing period. Full revenue platforms do not publish: Gong states that licences are per user, with a platform fee based on the number of users supported, and directs buyers to request a quote. Budget separately for the governance work, from field inventory, consent changes, CRM purge jobs, subject access requests and a contest path for reps — which is not on any pricing page and often exceeds the licence cost in year one.

In most deployments, no, and this is the single most common gap we would flag in an evaluation. Ask the vendor whether a rep can see every derived field about themselves, and whether there is a documented correction workflow. If the answer is that managers can share feedback, the answer is no. A score a person cannot see is a record they cannot correct, and in jurisdictions with data subject access rights it is also a disclosure obligation waiting to be exercised.

Inside the platform, usually yes — Gong's documentation, for example, states that deleting a call removes transcripts, media, notes and call stats. Outside the platform, almost certainly not. Any derived field already written to your CRM, posted into a chat channel, loaded into your warehouse or exported to a spreadsheet is governed by that system's retention rules, not the platform's. Deleting the call and believing the derived record is gone is the most common retention mistake in this category.

Choose on which decision the output feeds, not on the feature list. Sales-focused tools optimise for pipeline: deal health, forecast signals, competitor mentions, CRM write-back. Service-focused tools optimise for queue economics: resolution, escalation prediction, quality scoring across a large agent population. The governance profile differs too. Service deployments score far more employees far more often, which raises the stakes on the contest path, while sales deployments write more inferred fields into a system of record that finance and legal also read.

Run one quarter with one team and measure three things: how often a manager clicked through from a derived field to the actual call audio, how many derived fields nobody looked at, and how many tier-2 tags were wrong when spot-checked against the transcript. The first tells you whether the product is doing its real job of sending humans to the right thirty seconds. The second tells you what to switch off. The third gives you an error rate for the layer people wrongly assume is exact. Do not measure payback in the first quarter and do not let a vendor do it for you: the honest early result is a usage pattern and an error rate, not a revenue number, and no published methodology we could find supports the payback figures this category markets on.

Because the adoption problem is trust, not training. Reps stop engaging when a dashboard makes claims about them they cannot check or contest, and managers stop engaging when the scores contradict what they hear on calls they attended. Both failures trace to the same cause: tier-3 output presented with the same visual authority as tier-1 counts. Rollouts that survive tend to lead with search and shared listening, and introduce inferred fields last, if at all.

Split it deliberately. The revenue or service function owns the use case and the coaching process. IT and security own the connector scope, the tenant architecture and the audit trail. Someone must own the derived record itself — the field inventory, the retention clock in each destination system, and the contest path — and that ownership is what usually goes unassigned. Data that nobody owns is data nobody deletes, and derived data is the easiest kind for everyone to assume is somebody else's.

Ready to Govern Your AI?

Talk to LeapForce — one controlled layer for every AI tool, connector, model, and agent.

Thirty minutes · No pitch deck

Ready to turn AI experiments into measurable ROI?

Bring one outcome you'd like AI to move. We'll help you scope a pilot you can actually measure — and tell you honestly if it's not worth doing yet.

Comments