AI Customer Segmentation: The Point It Starts Deciding

AI customer segmentation is legally uneventful right up to the moment a segment starts deciding something. Grouping people to understand them is analysis. Group

AI customer segmentation is legally uneventful right up to the moment a segment starts deciding something. Grouping people to understand them is analysis. Grouping people so that the group determines the price they see, the offer they get, the credit limit they are allowed, or whether they ever reach a human, is an automated decision, and a different body of rules applies from that point on.

Our position is that the useful question is not which clustering method to use. It is where your decision line sits: the exact rung on which a segment stops describing a customer and starts deciding for them. Everything expensive about segmentation governance, from proxy discrimination to explainability to Article 22 to stale segments still firing eighteen months later, lives on one side of that line and is basically absent on the other. An underwriting modeller on Hacker News described the discipline that side of the line demands, writing that because "zip code is highly predictive of race, that attribute must also be excluded" (xeRTRex, Hacker News, November 2020). He was building insurance models, where that rule has been enforced for years. Most marketing teams building lookalike and value segments have never been told it applies to them at all.

The short answer: Classify every segment by what it is allowed to do — Describe, Suggest, Route, or Decide — and attach controls to the rung, not to the model. Only the top two rungs need Article 22 analysis, a proxy audit, an expiry date and a human who can overturn the outcome; treating all four the same either paralyses your analysts or leaves your pricing engine ungoverned.

Last updated: July 31, 2026.

Ladder diagram of four segmentation rungs from Describe to Decide with the controls each rung requires

The decision line sits between Suggest and Route. Below it, a segment informs a person. Above it, a segment moves without one.

We have not run a segmentation model against a real customer base of our own, and this article does not pretend otherwise. What we did do is run the proxy audit described below on a synthetic 20,000-customer dataset we generated ourselves, so that the numbers in the audit section are ones we produced rather than ones we quoted. They demonstrate the shape of the problem and what the audit output looks like. They prove nothing about anybody's real customers, and we say so again where the numbers appear. Everything legal below is quoted from the regulation, the case law or the regulator's own guidance, and linked, so you can check the wording rather than take our reading of it.

The decision line: where segmentation stops being analysis

The decision line is the point at which a segment's output changes something for a customer without a person choosing to change it. Below the line, segments are descriptions that humans read and act on. Above it, segments are instructions that systems execute. That single distinction predicts almost every compliance obligation, every audit request and every angry support ticket that segmentation generates.

Most segmentation guidance never draws it, because guidance is written from the model's point of view. From there, k-means over purchase history is the same operation whether the output goes into a slide deck or into a pricing API. From the customer's point of view the two are not remotely the same thing, and the law follows the customer.

The confusion is expensive in both directions. Teams that treat all segmentation as high-risk end up requiring an impact assessment before an analyst can plot a histogram, which is how you teach an organisation to route around its own governance. Teams that treat it all as harmless end up with a personalisation engine deciding who sees which price, with no owner, no audit trail and no way to answer the question a regulator asks first: why this person, in this group, on this day.

The second reason the line matters is the one that catches technical teams. Segments migrate upward. A segment built for a quarterly campaign readout gets wired into an email trigger, then a website module, then an offer engine, then a retention save-desk routing rule. Nobody decides to promote it. Each step is small and each is made by a different team. The definition never changes; what changes is what happens to you when you land in it. With no register of which rung each segment occupies, you cannot detect that drift, and by the time you do, the segment has been deciding things for a year.

The four rungs: Describe, Suggest, Route, Decide

Sort every segment you have into one of four rungs by asking one question: what happens to a customer purely because they are in this group? The answer places the segment, and the placement sets the controls. That is the whole framework, and it takes an afternoon to apply across a mature customer data platform.

RungWhat the segment doesReversible byArticle 22 exposureMinimum controls
1 · DescribeAppears in analysis, dashboards, research readoutsNobody needs to; nothing changedNoneNamed owner, documented definition
2 · SuggestSurfaces a recommendation a person chooses to act onThe person, before anything reaches the customerNone, if the person genuinely decidesOwner, definition, the person can see the reason
3 · RouteSends the customer down a different path — channel, queue, journey, service tierA person, but usually only after the customer complainsPossible, depending on effectEverything above, plus proxy audit, expiry date, logged assignment
4 · DecideSets price, offer, limit, eligibility, or access without a personOnly by escalation or appealLikely, if effects are significantEverything above, plus Article 22 analysis, human-intervention route, contestability, explanation on request

The decision line runs between rung 2 and rung 3. Rungs 1 and 2 are analysis. Rungs 3 and 4 are decisions, and everything difficult in this article applies to them.

Choose rung 1 if the segment exists to answer a question rather than trigger an action, and you can say honestly that no downstream system reads it. Keep it there by making the segment available in the warehouse and nowhere else. Most segments should live here and most companies would be better off if more of theirs did.

Choose rung 2 if a human is genuinely making the call. This is the underrated rung. A recommendation that a salesperson can ignore, a churn-risk flag a success manager reviews, a suggested discount band a pricing analyst approves — these get almost all of the commercial benefit at a fraction of the governance cost, provided the human involvement is real. The Article 29 Working Party guidance on this point is blunt about what "real" means: a controller "cannot avoid the Article 22 provisions by fabricating human involvement", and oversight must be "meaningful, rather than just a token gesture" carried out by someone "who has the authority and competence to change the decision" (Guidelines on Automated individual decision-making and Profiling, WP251rev.01). A queue of 400 approvals a day with an approve-all button is rung 4 wearing a rung 2 costume.

Choose rung 3 if the segment changes the customer's path but not their entitlement: a different onboarding journey, a different support queue, a self-service tier instead of a named account manager. Whether that is a "similarly significant" effect depends on what the different path costs them, which is why rung 3 needs the audit and the expiry date even when it probably sits outside Article 22.

Choose rung 4 only when you can carry the load. Price, credit, eligibility, access. If you cannot name the owner, produce the assignment log, run the proxy audit, and give a customer a human who can overturn the outcome, you are not ready to put a segment on this rung, and the honest move is to demote it to rung 2 with a person in the loop until you are.

What GDPR Article 22 actually says, and three things it does not

Article 22(1) of the GDPR gives a person "the right not to be subject to a decision based solely on automated processing, including profiling, which produces legal effects concerning him or her or similarly significantly affects him or her" (Article 22 GDPR). Three conditions have to stack up before it bites: the processing is solely automated, it produces a decision, and the decision has legal or similarly significant effects. Miss any one and Article 22 is not your problem. Hit all three and it is the centre of your compliance work.

That much is uncontroversial. What is routinely wrong in the vendor literature is the framing around it, and there are three specific errors worth naming because each one leads a team to the wrong build.

It is not a right you have to invoke. The most common misdescription treats Article 22 as an opt-out — something a customer must exercise, like unsubscribing. The Court of Justice of the European Union settled this in the SCHUFA judgment. Article 22(1), the Court held, "lays down a prohibition in principle, the infringement of which does not need to be invoked individually by such a person" (Case C-634/21, paragraph 52). The default is that the decision may not be made at all. You need one of the three gateways in Article 22(2), whether contractual necessity, Union or Member State law, or explicit consent — before you make it, not after somebody objects.

It is not limited to the system that takes the final action. SCHUFA also disposed of the argument that a scoring provider is merely supplying an input while the bank makes the decision. The Court ruled that automatically establishing a probability value about someone's ability to meet future payments is itself automated individual decision-making "where a third party, to which that probability value is transmitted, draws strongly on that probability value" to establish, implement or terminate a contract. The Court's reasoning was factual rather than doctrinal: an insufficient value "leads, in almost all cases, to the refusal of that bank to grant the loan applied for", so the score plays "a determining role". If your segment score is handed to a downstream team that follows it in almost every case, the score is the decision, whoever owns the button.

Consent is not the easy gateway. Article 22(2)(c) does allow explicit consent, and it is the option most often reached for because it feels like a checkbox problem. But Article 22(3) then requires the controller to implement suitable measures including, at minimum, "the right to obtain human intervention on the part of the controller, to express his or her point of view and to contest the decision". Consent buys you the right to make the decision; it does not remove the obligation to staff an appeals route. And Article 22(4) blocks the whole path where special-category data is involved unless narrow exceptions apply. A consent banner without a contest mechanism behind it is worse than no gateway at all, because it documents that you knew the rule applied.

Element of Article 22What it means for a segmentCommon misreading
"solely automated"No meaningful human in the loopA human clicking approve counts regardless of authority
"a decision"Includes a score a downstream party followsOnly the final action counts
"legal or similarly significant effects"Serious, sustained impact on circumstances or choicesAny personalisation qualifies
Prohibition in principleUnlawful by default without a gatewayA right customers must exercise
Article 22(3) safeguardsHuman intervention, view, contestConsent alone is sufficient

When a marketing segment crosses into Article 22 territory

Most marketing segmentation does not trigger Article 22, and pretending otherwise damages the credibility of the people who have to explain when it does. The regulator's own guidance says so plainly: "In many typical cases the decision to present targeted advertising based on profiling will not have a similarly significant effect on individuals", giving the example of a fashion advertisement aimed at "women in the Brussels region aged between 25 and 35" (WP251rev.01). If your segment decides which of two creative treatments a person sees for the same product at the same price, you are almost certainly fine.

The same guidance then lists exactly what flips it, and this list is the most useful compliance artifact in the whole document because it is short enough to check a segment against in five minutes. The factors are the intrusiveness of the profiling process, including tracking across different websites, devices and services; the expectations and wishes of the individuals concerned; the way the advert is delivered; and using knowledge of the vulnerabilities of the data subjects targeted.

Two of those four are worth dwelling on because they map directly to how modern segmentation is built.

Cross-device, cross-site tracking is the intrusiveness factor, and it describes the default architecture of a customer data platform. A segment assembled from first-party purchase history sits differently from one assembled by resolving identity across a data broker's graph. The guidance also warns that processing with little impact generally "may in fact have a significant effect for certain groups of society, such as minority groups or vulnerable adults", offering the example of someone in financial difficulty repeatedly targeted with high-interest loan adverts. That is the vulnerability factor, and it is the one where an optimisation objective quietly does the damage: a model trained to maximise conversion will find financially stressed customers for a high-cost credit product because they convert, and nothing in the training loop knows to stop.

Then there is price. The guidance states that automated decision-making resulting in differential pricing based on personal data "could also have a significant effect if, for example, prohibitively high prices effectively bar someone from certain goods or services". Segment-driven pricing is not automatically an Article 22 decision, but it is the segmentation use case most likely to become one, and it is the one growing fastest. The US Federal Trade Commission's surveillance pricing study found that behaviours "ranging from mouse movements on a webpage to the type of products that consumers leave unpurchased in an online shopping cart can be tracked and used by retailers to tailor consumer pricing", and that the intermediaries it examined worked with at least 250 client companies (FTC, January 2025).

There is a separate, simpler obligation attached to personalised pricing in the EU that has nothing to do with Article 22 and is missed constantly. Directive (EU) 2019/2161 inserted a new disclosure duty into the Consumer Rights Directive: traders must inform the consumer "where applicable, that the price was personalised on the basis of automated decision-making" (Directive (EU) 2019/2161). The accompanying recital explains the reasoning: consumers "should therefore be clearly informed when the price presented to them is personalised" so they can weigh the risk. Two limits are worth stating precisely, because both are commonly dropped. The duty sits in Article 6(1) of the Consumer Rights Directive, which governs distance and off-premises contracts rather than every transaction. And the recital expressly excludes dynamic or real-time pricing that carries no personalisation element, so a fare that moves with demand for everybody is not caught. Within those bounds it is a labelling requirement, it applies whether or not your pricing model is high-risk, and satisfying it costs a sentence of interface copy. It is the cheapest compliance win available to anyone doing segment-driven pricing in Europe, and the reason to do it first is that the disclosure forces an internal conversation about which segments actually move price, which is the register you needed anyway.

Proxy discrimination: what we found when we ran the audit

Proxy discrimination is what happens when a model reconstructs a protected attribute from features that are not protected. Nobody puts ethnicity in the feature set. The model finds postcode, device tier and category mix, and those three carry enough signal that the outcome tracks the attribute anyway. It is not a bug in the clustering algorithm. It is a property of the world the data came from.

The canonical published demonstration is not a marketing case at all, which is part of why it travels so well. Researchers examining an algorithm used to identify patients for extra care found that at the same risk score, "Black patients are considerably sicker than White patients". The cause was a label choice rather than a feature choice: the algorithm predicted healthcare costs as a stand-in for healthcare need, and because less money is spent on some patients for reasons that have nothing to do with how ill they are, the proxy imported the disparity wholesale. Correcting it "would increase the percentage of Black patients receiving additional help from 17.7 to 46.5%" (Obermeyer et al., Science, 2019). A commercial analogue is easy to write: predicted lifetime value is to customer worth what cost was to health need, an available number standing in for the thing you actually care about.

We wanted to know what this looks like in a segmentation setting specifically, and what the audit costs to run. So we built one.

What we ran. We generated a synthetic base of 20,000 customers. A binary protected attribute was assigned to 30% of them and then never used again — it appears in no model, no feature set and no training label anywhere in the code. Customers were placed in one of 40 postcode districts, with the protected group concentrated in the less affluent districts, which is the ordinary residential pattern the audit exists to detect. Device tier followed affluence. Twelve-month spend was generated from affluence, engagement and noise. We then trained a gradient-boosted regressor to predict spend from six commercial features — district, device tier, tenure, session count, returns flag and category mix — took the top 20% of predicted spend as a "high-value" segment, and measured two things: the selection rate for each group, and whether a second model could recover the protected attribute from the same six features.

This is synthetic data. It says nothing about any real customer base. What it demonstrates is the mechanism and the audit output. The script is about seventy lines, so the same two measurements can be put over a real segment in an afternoon by anyone who already has the data. Your numbers will not be these numbers: they depend on your population, and even on this synthetic base they shift with the random seed. Across thirty seeds, varying every random component, the selection ratio ran 0.23 to 0.32, proxy-recovery AUC 0.731 to 0.751, and the fit cost of dropping postcode 42.3% to 46.5%. The table reports one seed; that spread is its error bar.

Feature setModel fit (R²)Selected, protected groupSelected, other groupSelection ratioProxy recovery (AUC)
All six features0.7726.8%25.7%0.260.739
Postcode removed0.44111.3%23.8%0.470.643
Postcode and device removed0.38111.0%23.9%0.460.636

Read the first row carefully, because it is the result that matters. A segment built with no protected attribute anywhere in it selected one group at roughly a quarter of the rate of the other. And a classifier handed only the segment's own six commercial features recovered the protected attribute with an AUC of 0.739. That is meaningfully better than the 0.5 you would get from guessing. The segment did not merely correlate with the attribute; it encoded enough of it to be used as a substitute for it.

Neither number is visible from inside the segmentation workflow. Neither is expensive to produce, and that gap between cost and visibility is the whole reason this goes unmeasured.

Bar chart comparing model fit and selection ratio across three feature sets in the proxy audit

Removing the obvious proxy costs 43% of the model's fit and still fails the four-fifths threshold.

Why deleting the postcode did not fix it

The instinctive remedy is to delete the offending feature. Our audit says that instinct is half right and expensive, and the second and third rows of that table are the reason to publish the run at all.

Dropping postcode moved the selection ratio from 0.26 to 0.47. That is a real improvement and it is still a clear failure against any reasonable threshold. Dropping device tier as well moved it to 0.46, which is not an improvement at all. Meanwhile the model's fit fell from 0.772 to 0.441 and then to 0.381 — the first deletion cost about 43% of the model's explanatory power and bought a partial fix, and the second cost more and bought nothing.

The reason is that in the synthetic world we built, and in most real ones, the signal is not sitting in one column. It is distributed. Session count depends on affluence. Category mix depends on affluence. Affluence is what the geography encodes. Delete the postcode and the model routes around it through the behaviour, because the behaviour was generated by the same underlying structure. This is exactly the trap the Hacker News underwriting comment describes from the regulated side. The industry response there was not simply to delete zip code but to be able to "demonstrate that a model isn't discriminating", which is a testing obligation, not a feature-list obligation.

Three conclusions follow, and they are the opposite of what a feature-blacklist policy implies.

Blindness is not fairness, and a feature blacklist is not a control. Removing a column changes what the model can see, not what the world looks like. A policy that says "we do not use protected attributes" is a statement about inputs; the obligation is about outcomes. Recital 71 of the GDPR asks the controller to use "appropriate mathematical or statistical procedures" and implement measures that "prevent, inter alia, discriminatory effects" (Recital 71). Prevention of effects, not curation of inputs.

You need the protected attribute to measure the disparity. This is the paradox every team hits: to test whether your segment has adverse impact, you have to know the group membership you are forbidden to use. The resolution is architectural rather than legal. Hold the attribute in a separate, access-controlled store that the audit job can read and the production model provably cannot, log every read, and keep the two paths physically distinct. It is a governance design problem before it is a data science problem, and it is the single most common reason a proxy audit never gets run.

Fixing it costs accuracy, and someone has to sign off on how much. Our first deletion cost 43% of model fit. That is a commercial decision with a real number attached, and it belongs to a named person, not to whichever engineer was closest to the feature list. Write the number down. A fairness constraint applied quietly by a data scientist and a fairness constraint approved with its revenue impact stated are the same code and completely different governance.

The four-fifths check, borrowed from employment law

The threshold we used above is not a marketing standard, and we are not claiming it is a legal requirement outside its own domain. It comes from the US Uniform Guidelines on Employee Selection Procedures, which state that "a selection rate for any race, sex, or ethnic group which is less than four-fifths (4/5) (or eighty percent) of the rate for the group with the highest rate will generally be regarded by the Federal enforcement agencies as evidence of adverse impact" (29 CFR 1607.4(D), also on govinfo).

Borrow it anyway. It is a number, so it settles arguments that otherwise run on adjectives; it is decades old, so nobody has to be persuaded you invented a convenient threshold; and it is computable the moment you can join a membership list to a group attribute.

Borrow it honestly, though. The same regulation immediately qualifies itself: "smaller differences in selection rate may nevertheless constitute adverse impact, where they are significant in both statistical and practical terms", and greater differences may not, "where the differences are based on small numbers and are not statistically significant". So 0.82 is not a pass and 0.78 is not a violation. The ratio is a tripwire telling you where to look, and its real value is longitudinal. Run it monthly on every rung 3 and rung 4 segment. A ratio that was 0.91 in January and 0.66 in June is a story about something that changed in your data, and you want to find out what before somebody else does.

Two operational notes. Compute the ratio on the segment as deployed, not on the training set: the population that shows up in production is not the one the model was fitted on, and the gap is where surprises live. And run it separately for each decision the segment feeds. The same high-value segment can sit well within tolerance when it selects who gets an early-access email and badly outside it when it selects who gets a retention discount, because the base populations differ.

Explaining a segment to the person standing in it

Here is where AI customer segmentation gets genuinely hard, and it is not a compliance-theatre problem. A customer asks why they were offered a worse deal than their neighbour. The honest engineering answer is that a gradient-boosted ensemble put them above a threshold on a predicted value score computed from 60 features, and no single feature explains it. That answer is true, useless to the customer, and not sufficient.

The Court of Justice addressed exactly this in February 2025. Under Article 15(1)(h) of the GDPR, a data subject may obtain "meaningful information about the logic involved" in automated decision-making. The Court held that this requires the controller to explain "the procedure and principles actually applied in order to use, by automated means, the personal data concerning that person with a view to obtaining a specific result" (Case C-203/22, Dun & Bradstreet Austria).

Both halves of that ruling matter, and the relief is real. The Court said the obligation does not require "the mere communication of a complex mathematical formula, such as an algorithm" or a "detailed description of all the steps in automated decision-making". You do not have to hand over your model. What you do have to do is describe things so that, in the Court's words, "the data subject can understand which of his or her personal data have been used in what way" — in a "concise, transparent, intelligible and easily accessible form".

The Court also closed a door companies had been leaning on. A national provision that excludes access as a rule wherever a trade secret would be compromised is precluded; where the controller believes disclosure would expose trade secrets or third-party data, it must put that information before "the competent supervisory authority or court" so the balance can be struck case by case. Trade secrecy is a reason for a supervised balancing exercise, not a reason to say no.

What satisfies this in practice, for a segment, is four things written in advance rather than assembled under deadline:

  1. The purpose. What this segment is used to decide, in one sentence a customer would recognise as describing their situation.
  2. The categories of data used. Not the feature list with column names, but the kinds: purchase history over the last 24 months, service contacts, delivery region, subscription tenure.
  3. The top drivers for this individual. A per-customer attribution: the three or four inputs that moved this person's score most, produced by the same pipeline that produced the score. This is the only item that requires engineering, and it is the one that makes the other three credible.
  4. The route to a human. Who reviews it, how to reach them, what they are empowered to change, and how long it takes.

If you cannot produce item 3 for a rung 4 segment, you do not have an explainability problem to solve later. You have a segment that should be demoted to rung 2 until you do. That is an unpopular sentence in a roadmap meeting, and it is much cheaper than the alternative sequence, which is discovering it during a regulator's information request.

Diagram of the segment lifecycle from definition through decision attachment, drift, review and expiry

A segment has a lifecycle. Without an expiry date, only the first two stages ever happen.

The segment that stopped being true

Segments are built from a snapshot and applied indefinitely. That asymmetry is the quietest failure in the discipline, because nothing breaks. The pipeline runs, the dashboard is green, the segment keeps firing, and the only symptom is that it describes a person who no longer exists.

Someone is tagged as price-sensitive during a period when they were between jobs. Someone is tagged as low-value because they bought once, and then bought steadily through a channel the model does not read. Someone is tagged high-risk because of a delivery dispute that was resolved in their favour eighteen months ago. In every case the customer's behaviour changed and the segment did not, and in every case the customer has no idea a stale label is why they keep getting the cheap treatment.

The GDPR hook is the accuracy principle, which is more demanding here than the data-minimisation arguments people usually reach for. Article 5(1)(d) requires personal data to be "accurate and, where necessary, kept up to date", with "every reasonable step" taken to erase or rectify inaccurate data "without delay" (Article 5 GDPR). An inferred segment label is personal data. A label that was correct in 2024 and is wrong now is inaccurate personal data, and "the model has not been retrained" is not a reasonable step.

Three controls handle this, and none of them is exotic.

Give every segment an expiry date at creation. Not a review reminder: an expiry, after which membership stops being honoured downstream unless the segment has been rebuilt. The expiry belongs in the register and the enforcement belongs in whatever serves membership, so that letting a segment lapse fails safe. The interval is a business judgment. A lifecycle-stage segment might hold for a quarter, a stated-preference segment for a year, an intent segment for days.

Log the assignment, not just the membership. Store when a customer entered a segment, on which model version, and which decisions fired while they were in it. Without this you cannot answer the two questions that matter after a complaint — why were they in it, and what did it do to them — and you cannot compute the four-fifths ratio historically either. Store the model version alongside, because "which model" is the first question in any serious review and the hardest to answer retrospectively.

Give the customer a route to correct it. Article 22(3) already requires a contest route for rung 4 segments. Extending a lightweight version to rung 3 is cheap and it is the best data quality investment available, because the person with the strongest incentive to fix a wrong label is the person wearing it. A customer telling you the segment is wrong is a free relabelling with ground truth attached.

Reuse: the marketing segment that ends up deciding credit

The most dangerous thing that happens to a segment is not that it is wrong. It is that it is borrowed. A model built to decide who receives a newsletter is reasonable for that purpose and completely unvalidated for deciding a payment term. Nothing in a customer data platform distinguishes the two uses. A segment is a list of identifiers, and lists get copied.

The purpose limitation principle is the direct hook. Article 5(1)(b) requires personal data to be "collected for specified, explicit and legitimate purposes and not further processed in a manner that is incompatible with those purposes". A segment derived for marketing and then fed into a credit or eligibility decision is the textbook incompatible reuse, and the fact that it happened through three intermediate copies makes it harder to spot, not lawful.

The EU AI Act adds a sharper edge for a specific pattern of reuse. Article 5(1)(c) prohibits AI systems that evaluate or classify people "over a certain period of time based on their social behaviour or known, inferred or predicted personal or personality characteristics" where the resulting social score leads to "detrimental or unfavourable treatment of certain natural persons or groups of persons in social contexts that are unrelated to the contexts in which the data was originally generated or collected", or treatment "that is unjustified or disproportionate to their social behaviour or its gravity" (EU AI Act, Article 5).

Read the first limb slowly, because the phrase doing the work is "in social contexts that are unrelated to the contexts in which the data was originally generated". That is a description of segment reuse. This provision is aimed at social scoring, and we are not claiming that a retailer's high-value segment is a social credit system — that reading would be alarmist and would not survive contact with a lawyer. What we are saying is that the drafters chose context transfer as the thing that makes a classification dangerous, which is the same instinct behind purpose limitation and a good instinct to design around regardless of whether the prohibition reaches you. We have written more about how the AI Act's obligations land on companies deploying rather than building models in our guide to EU AI Act compliance for AI deployers.

The practical control is a purpose field on the segment that downstream systems must match. A segment carries the decision classes it is approved for; a system requesting membership declares what it is going to do with it; a mismatch is refused and the refusal is logged. This is boring plumbing and it is the only thing that actually stops reuse, because policy documents do not travel with data and access rules do. The related failure — the same customer record copied into a second store where the original controls do not apply — is one we have written about separately in the second-copy problem in AI workplace search.

The segment register: eight fields before the model ships

Everything above reduces to one artifact. If your organisation keeps a register of segments with these eight fields filled in, you can answer a regulator, a customer and an internal auditor without a project. If it does not, each of those requests becomes an archaeology exercise across a warehouse, a CDP and three campaign tools.

FieldWhy it existsFails without it
OwnerA named person, not a teamNobody can approve a change or answer a complaint
RungDescribe, Suggest, Route or DecideControls cannot be attached proportionately
PurposeThe decision classes this segment may feedReuse is undetectable
Definition and model versionWhat it means and what produced it"Which model decided this" is unanswerable
Data categoriesThe kinds of data used, in customer-facing languageArticle 15 access requests take weeks
ExpiryWhen membership stops being honouredStale segments fire indefinitely
Last auditDate and selection ratio per decisionDrift is invisible until it is a complaint
Contest routeWho can overturn an outcome and howArticle 22(3) is unsatisfied for rung 4

Two of these are missing almost everywhere, and they are the two that make the rest work: rung and purpose. Owner is often nominally present and functionally absent, a distribution list rather than a person. Expiry is almost never present at all.

The register is not a compliance document. It is an operations document that happens to satisfy compliance, and the difference shows up in whether anyone maintains it. Keep it where the segments are defined, make the rung field required, and make the serving layer read the expiry.

Choose your control point: model, segment, or decision

There are three places to put controls, they cost very different amounts, and most teams pick the most expensive one by default because it is the one data scientists can build alone.

Control pointWhat you enforceCostBest when
At the modelFairness constraints, feature exclusions, retraining disciplineHigh, per modelYou have few models, stable, owned by one team
At the segmentRung, purpose, expiry, audit cadence, registerModerate, onceYou have many segments across many tools
At the decisionWhich segments may feed which decisions, logged and refused on mismatchModerate, needs a chokepointDecisions are made by systems you can put a gate in front of

Choose the model control point if you have a small number of high-stakes models with clear ownership: a single pricing model, a single credit model. Fairness constraints applied in training are the strongest control available when they are feasible. They stop being feasible the moment you have forty segments produced by six tools, because you would be constraining forty things independently and someone will build the forty-first.

Choose the segment control point if you are a normal company. The register is where the leverage is: it is one artifact, it does not require retraining anything, and it survives model changes. This is the recommendation for most readers of this article, and it is deliberately unglamorous.

Choose the decision control point if you already have a place every automated decision passes through, or can create one. Enforcing at the decision is the strongest of the three because it catches reuse, catches expiry and catches new consumers of an old segment, none of which the other two do reliably. It is also the one that requires architecture you may not have.

When the incumbent still wins. If your segmentation is entirely rung 1 and rung 2 — analysis and human-reviewed recommendations — a spreadsheet register maintained by the analytics lead is genuinely sufficient, and buying tooling for it is waste. Adding a governance layer to a company with eleven descriptive segments and no automated decisions solves a problem it does not have. The trigger to invest is not the number of segments; it is the first segment that reaches rung 3.

Where a governance layer fits, and where it does not

LeapForce does not build your segmentation model, does not run your fairness audit, and has no view on whether k-means or a gradient boosting model is right for your customer base. That work stays with your data team, and any vendor telling you a governance layer solves proxy discrimination is selling you something that does not exist.

What the layer does own is the part of this article that is about access, identity and record rather than about statistics. Once a segment reaches rung 3 or 4, the questions stop being modelling questions and become the four this platform is built around: who is allowed to query this segment, what is the automated consumer of it allowed to do, what did it actually do, and what did it refuse to do. LeapForce is one controlled layer for every AI tool, connector, model and agent. The AI Gateway puts a single governed endpoint in front of model calls so a policy check happens at the point of use rather than in a document; Access and Identity treats non-human identities as first-class, so the service that reads a segment has a named owner, a scope and an expiry the same way a person does; and Observability and Audit keeps a tamper-evident action record that includes what was refused, not only what ran, which is the half of the record a complaint investigation actually needs. Those are the platform's published positions rather than a claim about your environment, and LeapForce discloses per-capability build status honestly, so ask which parts are live today before you plan around any of them.

Our rollout method is the same one we apply to any gateway deployment: observe first, enforce second, optimize third. Applied to segmentation, observe means finding out which systems currently read which segments before you write a single policy. In our experience of governance rollouts that is where the surprises are, because the answer is usually more systems than the register says. of governance rollouts generally is where the surprises are, because the answer is always more systems than the register says.

Honest limits on this analysis

The audit numbers in this article come from data we generated. They demonstrate a mechanism and they show what the output of a proxy audit looks like; they are not evidence about any real customer base, any real disparity, or any real model's accuracy cost. If you run the same audit on your own data you should expect different numbers, and the direction of the effect is not guaranteed to match ours.

We are not lawyers and this is not legal advice. Every legal claim above is quoted from the primary text and linked so you can read it yourself, but the application of Article 22 to a specific segment is fact-dependent, varies by supervisory authority, and turns on details of your particular processing that no article can assess. Rung 3 in particular sits in genuinely contested territory. Whether routing a customer to a different service tier is a "similarly significant" effect is exactly the sort of question that gets resolved case by case, and we have deliberately described it as "possible, depending on effect" rather than pretending there is a settled answer.

The four-rung model is our framework, not a recognised standard. It is a way of organising controls proportionately; it has no legal status and a regulator will not accept "it was rung 2" as an answer. Its value is internal. It makes the conversation about a specific segment concrete, which is more than most segmentation governance achieves.

Three things we could not resolve. We could not find published, non-vendor data on how long a typical commercial customer segment stays in production before being rebuilt, which is the number that would make the staleness section quantitative rather than argumentative. We could not find a credible recorded conference talk on proxy discrimination in commercial segmentation specifically, so this article carries no video. And the fastest-moving area here is US state law on surveillance and personalised pricing, which changed several times during 2026; we have deliberately cited only the federal FTC study rather than summarising state statutes, because a summary written today would be wrong somewhere by the time you read it.

 FAQ

Frequently asked questions

Yes, in general. Segmentation is profiling, and profiling is lawful processing provided you have a legal basis, meet the transparency obligations, and respect the data subject's rights. What changes the analysis is not the segmentation but the decision attached to it. A segment used for analysis or to inform a human sits under the ordinary GDPR rules. A segment that produces a solely automated decision with legal or similarly significant effects engages Article 22, which prohibits that decision by default unless one of three gateways applies.

Usually not, and the regulator says so. The Article 29 Working Party guidance states that in many typical cases, presenting targeted advertising based on profiling will not have a similarly significant effect. Four factors can change that: the intrusiveness of the profiling, including tracking across sites and devices; the expectations of the individuals; how the advert is delivered; and using knowledge of the data subjects' vulnerabilities. Differential pricing that effectively bars someone from a product is specifically flagged as capable of crossing the threshold.

Proxy discrimination is when a segment reconstructs a protected attribute from features that are not themselves protected. Postcode, device type, purchase history, channel preference. No protected attribute is in the model, yet the outcome tracks it. It survives because the features came from a world in which the attribute correlates with geography and behaviour. It is detected by testing outcomes, not by inspecting the feature list.

Run two tests on the segment as deployed. First, the selection-rate ratio: divide the share of one group selected into the segment by the share of the highest-selected group, and treat anything below 0.8 as a signal to investigate, following the four-fifths convention from the US Uniform Guidelines on Employee Selection Procedures. Second, the proxy-recovery test: train a classifier to predict the protected attribute from the segment's own features and check the AUC. Materially above 0.5 means the feature set encodes the attribute. Run both per decision the segment feeds, and run them monthly so you see movement.

In the EU, yes, where the price was personalised on the basis of automated decision-making. Directive (EU) 2019/2161 added that disclosure to the Consumer Rights Directive's pre-contractual information duties, and the recital explains the purpose is so consumers can factor the risk into their decision. It is a labelling obligation and it is separate from, and much easier to satisfy than, anything Article 22 requires. Rules elsewhere differ and are changing; check your own jurisdiction.

Not by publishing the model. The Court of Justice held in Case C-203/22 that "meaningful information about the logic involved" means the procedure and principles actually applied, expressed so the person can understand which of their data was used in what way — explicitly not a mathematical formula or every step of the process. In practice that means four written items: the purpose the segment serves, the categories of data used in plain language, the top three or four drivers for that specific individual, and how to reach a human who can change the outcome.

There is no universal interval, and any article giving you one is guessing. Set it per segment from how fast the behaviour underneath it changes: intent signals decay in days, lifecycle stage in a quarter, stated preference in a year. The design choice that matters is not the number but the mechanism. Set an expiry at creation and make the serving layer refuse expired memberships, so that neglect produces a visible failure instead of a silently stale label. Article 5(1)(d) of the GDPR requires personal data to be accurate and kept up to date where necessary, and an inferred label is personal data.

Treat that as prohibited until specifically validated. It is the clearest case of incompatible further processing under Article 5(1)(b), and a model validated for one purpose carries no evidence about its behaviour in another. Beyond the legal point, the practical one is that a marketing segment was optimised for response, not for default risk, and the errors it makes are cheap in one context and expensive in the other. Enforce it with a purpose field the requesting system must match, and log every refusal.

No. It is necessary in most regulated contexts and it is not sufficient anywhere. In our synthetic audit, a model with no protected attribute in it selected one group at 0.26 times the rate of the other, and deleting the most obvious proxy moved that only to 0.47 while costing 43% of the model's fit. The obligation in Recital 71 of the GDPR is to prevent discriminatory effects, which is a statement about outcomes. You have to measure the outcome, which means holding the protected attribute somewhere the audit can read it and the production model cannot.

A named individual per segment, recorded in the register, with authority to change or retire it. Ownership usually splits: the data team owns the model and its accuracy, the business function owns the purpose and the decisions attached, and one accountable name spans both. What fails is ownership by a distribution list, because complaints and audit requests need someone who can answer within a day. At rung 3 or above, that owner should also sign off the accuracy cost of any fairness constraint.

It does not regulate segmentation as such. The provision closest to it is Article 5(1)(c), which prohibits AI systems that classify people based on social behaviour or inferred personal characteristics where the resulting score leads to detrimental treatment in social contexts unrelated to where the data was generated, or treatment disproportionate to the behaviour. That is aimed at social scoring rather than at commercial segments, and ordinary marketing segmentation is not what it targets. The design lesson worth taking is the one about context transfer, which is the same principle as purpose limitation.

Only if the human involvement is real. The Article 29 Working Party guidance says a controller cannot avoid Article 22 by fabricating human involvement, that oversight must be meaningful rather than a token gesture, and that it must be carried out by someone with the authority and competence to change the decision, considering the relevant data. A reviewer processing hundreds of cases a day with no practical ability to reach a different conclusion is not human involvement, and describing the process that way in a data protection impact assessment documents the problem rather than solving it.

Ready to Govern Your AI?

Talk to LeapForce — one controlled layer for every AI tool, connector, model, and agent.

Thirty minutes · No pitch deck

Ready to turn AI experiments into measurable ROI?

Bring one outcome you'd like AI to move. We'll help you scope a pilot you can actually measure — and tell you honestly if it's not worth doing yet.

Comments