Product · Model Routing
The right model for every task — at the right cost.
Route each task on cost, speed, privacy, and quality instead of sending everything to the most expensive model. Set dollar budgets per team and agent, get alerts before the invoice does, and keep sensitive data on models that never leave your network.
Why routing
Most AI bills are paying frontier prices for commodity work.
40–60%
of production token spend is typically avoidable — redundant calls, over-specified models, and no caps.
Spend is growing faster than the teams driving it.
Industry trackers describe model spend doubling within months as agentic workloads multiply token consumption. The uncomfortable finding underneath: audits routinely attribute the majority of cost to a small set of high-volume, low-complexity tasks running on models dramatically stronger — and pricier — than the task requires.
The remedy isn't “use cheaper models.”
It's using the cheapest model that clears the quality bar for each task, and proving the bar is cleared. Routing research has shown that well-designed routers retain nearly all of frontier-level quality on mixed workloads at a fraction of the cost — precisely because most requests never needed the frontier.
Routing is also a privacy instrument.
The strongest guarantee for sensitive data isn't a clause — it's placement. A local model inside your network is one that external routes simply never see, so the route itself becomes the control.
How it works
Four routing strategies, combined per workload.
Match the model to the task
Requests are classified by difficulty: routine classification and drafting go to low-cost models; multi-step reasoning earns a premium model. The highest-leverage strategy for most companies.
Route by data class
Data tagged sensitive is pinned to local or in-tenancy models. This rule outranks every other — cost never negotiates with privacy.
Escalate on failure
Try the inexpensive model first; if the output misses a quality check, the request escalates a tier. You pay the premium only for the calls that need it.
Fail over, stay up
Provider outage or rate limit? Traffic shifts to the fallback chain automatically. Your teams notice a news story, not a work stoppage.
Quality is the guardrail on all of it: routing changes ship with evaluation gates checked against real task outcomes, so the cheap-model share rises on evidence — never on hope. Savings that quietly degrade answers aren't savings; they're deferred churn.
Budgets & attribution
Budgets in dollars, not tokens — enforced before the bill.
Token limits lie: the same ten thousand tokens cost pennies on one model and real money on another. Leapforce meters in currency, hierarchically — company, team, coworker, and tool each carry their own cap and pacing alerts. A looping agent hits its ceiling and stops; a team pacing hot gets a warning mid-month, not a surprise at month-end.
Every unit of spend is attributed — user, team, agent, task — so finance can charge back to cost centers your ERP understands, and leaders can see cost per task trend down as routing rules improve.
Live spend, metered in currency
Capabilities in depth
Everything between “one model for everything” and chaos.
Dynamic model selection
Routing rules pick the best model per agent and per task across every connected provider. When a better or cheaper model ships, adopting it is a rule change — not a quarter of migrations.
Evaluation gates
Routing changes are validated against task-level quality checks before and after rollout. If the cheap tier degrades outcomes for a workload, the rule rolls back — visibly.
Local model option
Open-weight models run on your GPUs — in your cloud tenancy or on-prem — and receive the workloads policy pins to them. Privacy by placement, with the audit trail to show it.
Pacing alerts & chargeback
Thresholds notify owners before caps hit; monthly attribution exports map spend to cost centers. The AI bill becomes a managed line item with names on it.
Keep existing subscriptions
Teams keep the AI services they already trust; the company manages access and spend centrally. Duplicate seats and orphaned subscriptions surface in the same view.
Per-environment rules
Development and testing default to inexpensive tiers; production earns the stronger models. CI pipelines and experiments stop competing with customer-facing work for premium capacity.
Waste elimination
Response caching for repeated questions, batch execution for latency-tolerant jobs at provider discounts, and prompt hygiene that trims token-inflated inputs — the unglamorous savings that add up.
Agent-aware accounting
Agentic runs consume tokens across many calls; routing tracks cost per run and per outcome, so “what does a resolved ticket cost us?” has a number that improves over time.
Questions we hear
Model Routing — frequently asked.
Won’t cheaper models make our answers worse?
Not if routing is governed by evaluation. Most enterprise volume is genuinely simple — that’s why research routers preserve nearly all frontier quality on mixed workloads. Every routing rule here ships with a quality gate on real task outcomes, and rules that degrade results roll back. Cost falls where quality holds; nowhere else.
What happens when a budget cap is hit?
Configurable per scope: hard-stop (requests refuse politely), downgrade (traffic shifts to a cheaper tier), or notify-and-continue for critical workloads. Owners are alerted at pacing thresholds well before the cap.
How much will we actually save?
It depends on your mix, so we won’t quote a universal number. Industry field audits regularly find a large share of token spend avoidable; your first weeks on the gateway produce the attribution data that turns that generic claim into your specific answer — before any routing rule changes a thing.
Do we need our own GPUs for local models?
Only for the local route. It runs on GPU capacity in your cloud tenancy or on-prem hardware, sized to the sensitive workloads you pin there — most companies start small, because most volume doesn’t need it.
Who decides the routing rules?
You do, with defaults to start from. IT and finance set budget hierarchies; workload owners tune model tiers per task with the evaluation data in front of them. On the managed plan, we run the tuning cycle with you.
How do provider price changes affect us?
Pricing shifts show up in the attribution data immediately, and rules can respond — shifting tiers or providers — without any application changes. Decoupling apps from providers is half the point of routing through a gateway.
Can a team insist on a specific model?
Yes — pinning a workload to a named model is a legitimate rule, and its cost is attributed like any other. Routing is about making the trade-off explicit, not taking the choice away.
Is routing a separate product to buy?
It’s part of the platform — routing runs inside the AI Gateway, using the same identity, policy, and audit plane. Plan specifics are covered on a strategy call.
Leapforce is in active development — per-capability build status (live, in development, roadmap) is disclosed honestly on request. Industry figures above are drawn from public 2025–2026 research and cited ranges vary by study.