Devin pricing in 2026 is Free at $0, Pro at $20/month, Max at $200/month, Teams at $80/month plus $40/month per full dev seat, and Enterprise by quote — all fetched from Devin's pricing page on July 31, 2026. Those are the small numbers. The large one is the engineer who reads the output before it merges.
Our position is that Devin's meter is the wrong thing to negotiate. Cognition's own documentation says it plainly: "the bottleneck shifts from writing code to reviewing it" (Devin Review docs, fetched July 31, 2026). Once an agent opens pull requests faster than a team can read them, the invoice line that grows is payroll, not compute. That is the shape of the complaint on the ground too. In an April 2026 Ask HN thread titled "Am I getting old, or is working with AI juniors becoming a nightmare?", the commenter decasteve described the daily reality: "I spend way too much time explaining AI generated code" to the developer who nominally wrote it. He was not complaining about a bill. He was describing one.
The short answer: Price Devin as (metered spend + review hours × loaded hourly cost) ÷ merge rate. On every published number we could find, the meter is the smallest term in that expression and the merge rate is the one that decides whether the deal works.
Last updated: July 31, 2026.

Metered spend, review hours, and the merge rate that divides both.
We have not run Devin ourselves, we have not measured a merge rate, and nothing below is a report of our own usage. Every figure here is either a number Cognition publishes, a number a named third party published, or arithmetic applied to those numbers with the assumptions stated in the open. Where we model, we say we are modelling.
Devin pricing in 2026: every number Cognition publishes
There are five Devin plans. Four are self-serve and one is quoted. Fetched from devin.ai/pricing on July 31, 2026, the ladder runs Free at $0, Pro at $20 per month, Max at $200 per month, Teams at $80 per month plus $40 per month per full dev seat, and Enterprise at "Let's talk."
| Plan | Published price | Members | Concurrent sessions | The thing that actually gates you |
|---|---|---|---|---|
| Free | $0 | 1 | Up to 10 | Light quota, limited model availability |
| Pro | $20/month | 1 | Up to 10 | Single user; cannot invite anyone |
| Max | $200/month | 1 | Unlimited | Still a single user, just a bigger quota |
| Teams | $80/month + $40/month per full seat | Unlimited | Unlimited | $80 floor whether or not you use it |
| Enterprise | Quoted | Unlimited | Not published | ACU volume set in your order form |
Three details on that table cost more than they look.
Pro and Max are individual plans. Devin's self-serve billing documentation, fetched July 31, 2026, is unambiguous: "Pro subscribers cannot invite additional members to their organization." Max is the same. So the second person who needs an account moves you to Teams, and Teams carries an $80 monthly minimum regardless of how little you consume. That is a seat cliff, not a volume cliff. It arrives at headcount two.
Teams then splits members into full seats and flex seats. A full seat costs $40 per month, includes a Pro-equivalent daily and weekly quota, and grants access to Devin Desktop. A flex seat is free, has no fixed charge, excludes Devin Desktop, and draws entirely from the team's shared pool of on-demand credits. Unlimited flex seats are permitted.
The $80 floor is met in whatever combination you land on, and the docs publish the table:
| Full seats | On-demand credits included | Monthly total |
|---|---|---|
| 0 | $80 | $80 |
| 1 | $40 | $80 |
| 2 | $0 | $80 |
| 3 | $0 | $120 |
| N ≥ 2 | $0 | N × $40 |
Below two full seats, the remainder of the minimum is charged as prepaid on-demand credits the whole team can spend. Those credits, per the same page, "roll over month-to-month," and "purchased credits never expire" — a genuinely friendly term, and worth naming, because prepaid credit that expires is the standard trap in this category and Devin has not set it.
Enterprise is where the second unit appears. Devin's enterprise billing page states that Enterprise customers "are billed in Agent Compute Units (ACUs) at the rate set in their order form" and directs you to sales for pricing. There is no published Enterprise rate card. What is published is the control: enterprise admins can set per-Organization ACU limits, and "all Devin activity stops once the limit is reached."
The unit you cannot price: what an ACU is and what Devin stopped publishing
An ACU is Devin's Agent Compute Unit: a measure of work performed, not of time elapsed. Per Devin's usage documentation, fetched July 31, 2026, usage accrues on "number and complexity of actions Devin takes (planning, context gathering, task execution, browser actions, code execution, and so on)" plus virtual machine time and networking bandwidth, which the page describes as "typically a small fraction of total usage."
The same page carries three mechanics that matter for forecasting. Windows sessions consume roughly 9% more than equivalent Linux sessions. Aside from the few units needed to keep the VM running, Devin does not consume usage while sleeping, while waiting for your response, while waiting for a test suite, or while cloning repositories. And a session sleeps automatically "after roughly 0.1 ACUs" of inactivity, so an abandoned tab is not a runaway meter.
Here is the part almost nobody has updated. As of the pricing page we fetched on July 31, 2026, Devin publishes no self-serve dollars-per-ACU rate, no quota size, and no numeric overage rate. The usage page states that Enterprise customers consume ACUs against their order-form volume while self-serve customers "consume their plan's included quota first, then draw from prepaid on-demand credits." The pricing page's own FAQ says each paid plan "comes with a usage allowance that refreshes automatically on a daily and weekly basis" and that overage "is consumed at API pricing." Neither the allowance nor the API price is stated as a number.
Search results for "Devin pricing" are still dense with a three-tier ladder — Core at $20 with pay-as-you-go ACUs at $2.25, Team at $500 including 250 ACUs at $2.00, Enterprise by quote. Cognition itself now describes that structure in the past tense. The self-serve documentation carries a section headed "Migrating from legacy ACU-based plans" which states that "Legacy Core plan users have been migrated to the Free plan" and that "On-demand credits are the same dollar value as the ACUs you're used to." So the plan names in those articles are documented as legacy, and the one bridge Cognition does publish is an equivalence — a credit is worth what an ACU was worth — with no dollar figure attached to either side. Treat any per-ACU dollar figure you read in a third-party article as historical unless it sits on a dated Cognition page. We are not asserting the old rates were never real. They are simply not what the vendor sells today, and a procurement model built on them is pricing a product that has been retired.
That has a direct consequence for buyers. You cannot forecast the meter from published data, because the vendor does not publish its price. What you can forecast, from your own systems, is the review burden. That asymmetry is the whole argument of this article.
There is one published number that gives ACUs a rough shape. Session Insights classifies every completed session by size, using ACU count and user-message count together:
| Session size | ACU threshold | User message threshold |
|---|---|---|
| XS | ≤ 2 ACUs | ≤ 2 messages |
| S | ≤ 5 ACUs | ≤ 5 messages |
| M | ≤ 10 ACUs | ≤ 10 messages |
| L | ≤ 20 ACUs | ≤ 20 messages |
| XL | > 20 ACUs | > 20 messages |
The overall size is the larger of the two classifications, enterprise ACU thresholds are scaled by a factor of ten, and sessions landing at L or XL are "flagged as unhealthy, meaning Devin likely encountered significant issues or the task scope was too broad for a single session."
Read that table as a review-cost table rather than a compute table. A session that took twenty of your messages to steer is a session where a human was already in the loop for the whole run — before the pull request even existed to be reviewed.
Where the meter ends and the review hour begins
The review hour is not a governance abstraction; it is the load-bearing cost in agentic coding, and three published sources point the same direction. Cognition's own 2025 performance review, published November 14, 2025, states that "Only 20% of engineering time is spent coding; much more goes into other work, like planning and reviewing." An agent that compresses the 20% while adding to the 80% is not automatically a saving.
The same post is candid about where the human stays. On first-pass code review it says Devin "can execute first-pass reviews and catch obvious issues," then immediately: "Human review is still necessary, because code quality is not straightforwardly verifiable." That sentence is the vendor telling you the review hour does not go away.
Two independent findings sharpen it. DORA's 2025 State of AI-assisted Software Development report, based on nearly 5,000 technology professionals, found 90% of respondents using AI at work and more than 80% reporting a productivity increase — alongside 30% reporting little or no trust in AI-generated code, and a continued negative relationship between AI adoption and software delivery stability even as throughput improved. Faster in, less stable out, is exactly the profile that converts into review and rework.
The second is the one that should make any buyer slow down. METR's randomised controlled trial, published July 10, 2025, had 16 experienced open-source developers complete 246 issues on their own repositories, with AI tools allowed on a random half. Developers expected AI to speed them up by 24%. They were measured taking 19% longer with AI. After the study, having lived through the slowdown, they still believed AI had sped them up by 20%.
That gap is the reason this article exists. If practitioners misjudge their own throughput by roughly forty points in the optimistic direction, then a buying decision based on how much faster the team feels is not a costing exercise at all. The meter you can see is small; the cost you cannot see is the one people misestimate.
The counter-case is real and it is published, so we will not dodge it. Cognition reports customer outcomes that are large and specific: one large organisation "saved 5-10% of total developer time" using Devin for security fixes; another moved vulnerability handling from "30 minutes per vulnerability" for a human developer to 1.5 minutes for Devin; a bank migrating proprietary ETL framework files saw each file completed in "3-4 hours vs 30-40 for human engineers." Nubank's published case study claims 8-12x engineering time efficiency and over 20x cost savings on the portion of its ETL migration delegated to Devin, against a codebase of more than six million lines and roughly 100,000 data class implementations. These are vendor-published customer numbers, not independent audits, and we have not verified them against Nubank's own reporting — but they are consistent in shape, and the shape is instructive: every one of them is a high-volume, pattern-repetitive, machine-verifiable task. That is the regime where review per unit of output collapses, because the reviewer checks a pattern once and a diff a thousand times.
Here is a practitioner conversation on measuring exactly this — Amos Haviv, who leads Developer Workflow teams at Booking.com across roughly 4,000 engineers and 8,000 repositories, on how to prove AI is actually shipping more features rather than feeling like it:

The Merged-Work Cost model
Merged-Work Cost is the framework this article proposes, and it fits on one line:
Cost per merged pull request = (M + R × H) ÷ m
Where M is metered spend per attempt, R is human review hours per attempt, H is the loaded hourly cost of the reviewer, and m is the fraction of attempts that merge.
Four properties make it worth using instead of a per-seat comparison.
It charges review to every attempt rather than to every merge. A pull request that gets read and rejected consumed a reviewer's time in full. Dividing by m at the end is what puts the cost of the rejected work back where it belongs — on the merged work that survived.
It makes the meter's true share visible. M appears once and unmultiplied. R × H is a person's hour. Unless a single agent attempt burns more compute than an engineer costs in the same window, the meter is a rounding error in this expression.
It has exactly one variable you can move quickly. H is set by payroll. M is set by a vendor who currently does not publish it. R falls slowly, as tooling and prompts improve. m is the term that swings by a factor of 4.5 between published data points, and it responds to task selection within a single sprint.
It is a forecasting model, not a measurement. We have not run these workloads. Every number in the next two sections is arithmetic on stated assumptions, and the assumptions are yours to replace.
The assumptions, stated plainly. For the loaded hourly cost H we use $110/hour. Derivation: the Stack Overflow 2025 Developer Survey reports a median total annual compensation of $175,000 for back-end developers in the United States (the global median for the same role is $79,742, so if you are not hiring in the US this number should fall a long way). Divided by 2,080 working hours that is $84.13. We then apply a 1.3x employer-load multiplier for benefits, payroll tax, equipment and facilities, giving $109.38, rounded to $110. The 1.3x is our chosen assumption, not a sourced figure — organisations use anything from 1.2 to 1.5 and you should substitute your own finance team's number.
For M we use $3 per attempt as an illustrative placeholder, because Devin does not publish a self-serve per-attempt price. We chose $3 deliberately as a generous figure: Pro costs $20 for a whole month, so if one attempt genuinely cost $3 of metered compute, seven attempts would exhaust the subscription. If your real figure is $1 or $10, the conclusions below barely move, which is itself the finding.
For R we model two cases: 30 minutes of review per attempt (a small, well-scoped, test-covered change) and 90 minutes (a bloated diff on unfamiliar code, the situation the Ask HN thread describes).
Three costed scenarios at three merge rates
Cognition publishes a merge rate. In the 2025 performance review it states that "67% of its PRs are now merged vs 34% last year." That is a vendor-reported figure across its own customer base, not an audited or independently replicated one, and your codebase is not the average of Cognition's customers. It is still the single most useful published number in this market, because it is the term the Merged-Work Cost model is most sensitive to.
At the other end sits the most rigorous public hands-on evaluation we could find. Answer.AI's write-up, published January 8, 2025, reports: "Out of 20 tasks we attempted, we saw 14 failures, 3 inconclusive results, and just 3 successes." That is a 15% success rate, over about a month, on an early version of the product. It is a small sample from a single team on a product that has changed substantially since. Cognition's reported merge rate doubled over that same period. Treat it as a floor case, not a current benchmark.
Modelling both ends, at 30 minutes of review per attempt, M = $3, H = $110:
| Merge rate | Source of the rate | Cost per attempt | Cost per merged PR | Meter share of per-attempt cost |
|---|---|---|---|---|
| 67% | Cognition, Nov 2025, vendor-reported | $58.00 | $86.57 | 5.2% |
| 50% | Midpoint scenario, ours | $58.00 | $116.00 | 5.2% |
| 15% | Answer.AI, Jan 2025, early version | $58.00 | $386.67 | 5.2% |
The per-attempt cost is identical in all three rows. The merge rate alone moves cost per merged pull request by a factor of 4.5. And across the whole table, metered compute is 5.2% of what an attempt costs you.
Now the review-hour case. Same rates, but 90 minutes of review per attempt:
| Merge rate | Cost per attempt | Cost per merged PR | Multiple vs the 30-minute case |
|---|---|---|---|
| 67% | $168.00 | $250.75 | 2.9x |
| 50% | $168.00 | $336.00 | 2.9x |
| 15% | $168.00 | $1,120.00 | 2.9x |
Tripling review time nearly triples the bill, at every merge rate. Tripling the meter, from $3 to $9, moves the 67% cell from $86.57 to $95.52 — a 10% change. That is the asymmetry in one comparison: a variable the vendor controls and does not publish is worth about a tenth of a variable your own engineering process controls entirely.
The break-even merge rate, and why it is lower than you think
The honest next question is not "is Devin expensive" but "expensive compared to what." Cognition defines its own sweet spot: Devin excels at "tasks with clear, upfront requirements and verifiable outcomes that would take a junior engineer 4-8 hrs of work." So take the alternative as the same task built by hand, and take the vendor's own range.
At H = $110/hour, a 4-hour task costs $440 to build by hand and an 8-hour task costs $880. Use the 6-hour midpoint, $660, as the benchmark B. Break-even is the merge rate at which Merged-Work Cost equals B:
m\* = (M + R × H) ÷ B
| Review time per attempt | Cost per attempt | Break-even merge rate vs a $660 hand-built task |
|---|---|---|
| 15 minutes | $30.50 | 4.6% |
| 30 minutes | $58.00 | 8.8% |
| 60 minutes | $113.00 | 17.1% |
| 90 minutes | $168.00 | 25.5% |
| 120 minutes | $223.00 | 33.8% |
Break-even merge rate against a $660 hand-built task, as review time per attempt rises.
This is where the analysis turns, and it turns in Devin's favour. At half an hour of review per attempt, you need fewer than one in ten attempts to merge before the agent beats hand-building the task. Against Cognition's reported 67%, that is an enormous margin — and even Answer.AI's 15% floor clears the 30-minute row. The economics of a coding agent are extremely forgiving of failure, as long as failures are cheap to identify.
The word doing the work in that sentence is identify. Every row in that table is driven by R, and R is not a property of the agent. It is a property of your codebase, your test suite, your diff size, and whether the reviewer can tell in ten minutes that an attempt is wrong or has to spend an hour finding out. Push review to two hours per attempt and you need a third of everything to merge just to break even — and now you are betting the whole business case on the one number the vendor reports about itself.
Two more things the table quietly assumes, and both are optimistic. It assumes a failed attempt costs review time but nothing else, when in practice a plausible-but-wrong change that gets merged costs far more than one that gets rejected — DORA's stability finding is that risk showing up in the aggregate. And it assumes reviewer time is available. In a team where two senior engineers are the only people who can approve changes to the payments service, agent throughput does not convert into shipped work at all; it converts into queue.
What actually moves the merge rate, according to Devin's own docs
If the merge rate is the decisive variable, the practical question is what raises it. Cognition's When to Use Devin guidance, fetched July 31, 2026, is unusually specific, and every item on it is a review-cost lever as much as a success lever.
| Lever, as published by Cognition | Why it moves cost, not just quality |
|---|---|
| "if a task would take you three hours or less, Devin can most likely do it" | Sets the unit of delegation. A three-hour task produces a diff a reviewer can hold in their head. |
| Tasks with "test suites, CI checks, or verifiable outcomes" | Moves verification from a human reading code to a machine running it. This is the single biggest reduction in R available. |
| Keep sessions at XS, S or M in Session Insights | An L or XL session is flagged unhealthy. It is also the session that produces the diff nobody wants to review. |
| Scope with Ask Devin before implementation | Front-loads the human judgement into a cheap step instead of an expensive one. |
| Devin Review with Auto-Fix, so Devin responds to review comments and CI failures | Removes the round-trip, not the review. The docs' own aspiration is "you just need to see that CI passes and the PR is approved." |
| Split large projects into sub-tasks across parallel sessions | Turns one unreviewable diff into several reviewable ones. |
Read as a group, that list says something the pricing page does not: the work that makes Devin cheap is work your team does, before the meter starts. Writing a good task description, having a test suite that actually fails when behaviour breaks, and keeping diffs small are all preconditions. None of them are billed by Cognition and all of them are paid for by you.
The corollary is uncomfortable for a certain kind of buyer. A team whose test coverage is weak, whose tickets are one-line hopes, and whose codebase has no clear patterns to imitate will get a low merge rate and a high review hour simultaneously — the two terms move together in the wrong direction. DORA's central 2025 finding was that AI "amplifies" existing strengths and weaknesses rather than fixing them. The Merged-Work Cost model is what that amplification looks like on an invoice.
Measure your own merge rate before you sign anything
You do not have to take anyone's merge rate on faith, including ours, and you do not need a Devin contract to start. The measurement is a two-week exercise against data you already have, and it produces the three numbers the model needs.
Week zero — establish the baseline you will be compared against. Pull the last 90 days of merged pull requests in the two or three repositories you would actually point an agent at. Record, per PR: lines changed, time from open to first review comment, time from open to merge, and number of review round-trips. You now have a human R and a human cycle time. Nobody can argue with your own git history.
Weeks one and two — run the trial as a costing exercise, not a demo. Devin's Free plan exists and Pro is $20, so the entry cost of measurement is trivial next to the cost of a wrong annual commitment. Delegate a fixed list of tasks chosen to match the vendor's stated sweet spot: clear requirements, verifiable outcomes, under three hours by hand. For each attempt record four things:
- Did it merge, merge after rework, or get abandoned? (This is m, and only the first two count.)
- Wall-clock minutes a human spent reading, testing and correcting it, including attempts that were abandoned. (This is R, and the abandoned ones are the ones teams forget to count.)
- The session's ACU cost and size classification, from Session Insights. (This is M's shape, and the size tells you whether the task was scoped right.)
- Whether the reviewer could have caught the failure from CI alone.
That fourth question is the one that predicts your second year. If most failures were catchable by a test, R collapses as you invest in tests. If most failures required a human to understand intent, R is structural and the model needs a merge rate well above the break-even row before it works.
The two-week exercise produces the four terms the model needs, from data you already hold.
End of week two — populate the model. You now have real values for m, R and M, an H from finance, and a B from your own baseline. Run the arithmetic once. If cost per merged pull request comes in under B, buy on the basis of that number and not the vendor's. If it comes in over, you have learned something worth far more than $20: which of the four terms is out of line, and whether it is fixable by scoping, by tests, or not at all.
We have not run this two-week exercise ourselves — it needs a real codebase, a real team and a real backlog, none of which we can borrow honestly. What we can tell you is that it uses only inputs a mid-sized engineering org already has, and that the first number it produces, the human baseline, is useful even if you never buy anything. This is the same discipline we argued for in our analysis of what enterprise AI actually costs beyond the licence, and the same one behind ranking agent types by the cost of oversight rather than by capability.
Seat math: where the plan structure bites
Devin's plan structure is cleaner than most in this category, but three edges catch buyers.
The individual-to-team cliff. Pro at $20 and Max at $200 are single-user plans that cannot invite members. The moment a second engineer needs their own account, the floor is $80. In practice, one full seat plus the $40 of credits that top up the minimum is the realistic two-person starting point — $80/month — and it climbs at $40 per additional full seat once you pass two.
Flex seats are a real option, and they are the right default for reviewers. A flex seat is free, has no monthly charge, and draws from the shared credit pool. It excludes Devin Desktop. For an engineer whose job in the loop is mostly to review Devin's output and occasionally kick off a task, that trade is close to free. Devin's own guidance is to reserve full seats for regular users and flex seats for occasional ones — worth taking literally, because a team that gives every engineer a full seat out of politeness is paying $40 a month per person for quota that a shared pool would have covered.
Enterprise buys you the ACU limit, and the limit is the actual product. Per-Organization ACU caps, set from Enterprise Settings, stop all Devin activity in that organisation when hit. That is a hard ceiling, which is more than most agent vendors offer, and it is the main structural reason a large buyer moves off Teams. Whether the negotiated per-ACU rate beats $40-per-seat-plus-credits is not something anyone can tell you from published data, because Cognition does not publish an Enterprise rate.
One planning note that follows from the meter's mechanics rather than the price list: because usage accrues on actions rather than elapsed time, and because sessions sleep after roughly 0.1 ACUs of inactivity, the thing that inflates a Devin bill is not people leaving tabs open. It is people running long, unscoped sessions with many corrective messages — which is the same behaviour that produces an XL session, an unreviewable diff, and a failed merge. The cost drivers and the quality drivers are the same drivers. That is unusual and it is good news.
Spend controls Devin gives you, and three it does not
Published controls, from the billing documentation fetched July 31, 2026:
| Control | Where it lives | What it actually does |
|---|---|---|
| Per-Organization ACU limits | Enterprise Settings > Organizations | Hard stop: all activity halts at the limit |
| Default session spending limits | Settings > Usage | Caps what a single session can consume |
| Auto-reload thresholds | Settings > Usage | Governs top-ups rather than restricting them |
| Shared on-demand credit pool | Teams | One balance the whole team draws from |
| Per-session ACU cost | Session Insights | Retrospective, per session, visible to any user |
| Consumption Analytics | Enterprise / Organization Settings | Breakdown by organisation and by user |
That is a better set than most agent platforms publish. Three gaps remain, and they are the ones a finance function will ask about.
There is no published dollar budget. Every control above is denominated in ACUs, credits or sessions. Translating a departmental budget in dollars into an ACU cap requires the Enterprise rate from your order form, which means the control is only as usable as your contract makes it. We argue elsewhere that budgets belong in dollars rather than tokens precisely because the unit a vendor meters in is rarely the unit a budget holder is accountable for.
The shared pool has no per-member balance. The self-serve documentation states plainly that on Teams, credits are shared across all members "with no per-member balance. Any teammate can draw from the shared pool." Combined with unlimited free flex seats, that means the number of people who can spend your balance is unbounded by design. The Enterprise tier answers this at the organisation level; Teams does not answer it at the person level.
Nothing meters the review hour. No control in the table above sees the cost this article is about. Session Insights will tell you an attempt cost 7 ACUs; nothing tells you it cost a staff engineer ninety minutes. That number lives in your issue tracker, your calendar, and your people's evenings, and if you want it you have to go and build it. This is the general pattern with agent platforms: the vendor instruments the half of the cost it bills for.
When the review-hour argument does not apply
This piece would be dishonest if it did not name the cases where the meter really is the main event and the review hour genuinely collapses.
Fleet migrations against a verified pattern. The Nubank shape — one repetitive transformation, applied across thousands of files, with a pattern a human validates once. Review cost per unit of output falls toward zero because the reviewer is checking conformance, not intent. Cognition's published customer numbers cluster in exactly this regime, and they are large.
Machine-verifiable defect classes. Vulnerabilities surfaced by static analysis, dependency and language-version upgrades, and lint or type errors all have an oracle. When CI can decide correctness, R is minutes and the break-even merge rate falls to the top row of the table.
Test generation with an existing playbook. Cognition reports coverage typically rising "from 50-60% to 80-90%," with humans checking logic afterwards. Reviewing a test is cheaper than reviewing a behaviour change, because a wrong test fails loudly and a wrong feature fails quietly.
Very high volumes on Enterprise terms. If you are consuming ACUs in the volumes those case studies imply, the negotiated rate genuinely becomes a real line item and deserves the negotiation. The argument in this article is about the median buyer running tens of tasks a month, not the buyer running hundreds of thousands of file migrations.
Conversely, the review hour dominates hardest on greenfield feature work with no established pattern, on changes to systems where correctness is contextual rather than testable, on codebases where the only qualified reviewers are already the bottleneck, and anywhere the failure mode is a plausible-looking change that passes CI and is wrong.
Where a governance layer fits, and where it does not
Nothing in this article is a LeapForce product problem until the agent stops being one engineer's experiment. At that point the questions change shape: which agent identity opened this pull request, what was it allowed to touch, who owns it when its author leaves, which budget did the run draw from in dollars, and can you produce the record twelve months later. Those are the questions LeapForce is built for: non-human identities with owner, scope and expiry; dollar-denominated budgets with chargeback rather than token or credit counters; tamper-evident records of what an agent did and what it was refused; and human approval gates as a first-class step rather than a convention. Our rollout model for the gateway is deliberately slow in the same way this article recommends: Observe first. Enforce second. Optimize third. — you cannot cap what you have not yet measured.
To be clear about the boundary: LeapForce does not review code, does not write pull requests, and is not an alternative to Devin. If your problem is that a coding agent's output takes too long to read, a governance layer will not read it for you. What it will do is make the spend attributable and the actions provable once several agents are running across several teams — That is a different problem. It usually arrives some months after this one.
Honest limits on this analysis
The meter is genuinely unpriced for self-serve, so a whole class of question is unanswerable here. We could not find a published dollars-per-ACU figure, a published quota size, or a published API overage rate on any Cognition property on July 31, 2026. The closest thing published is the legacy-migration note stating that on-demand credits carry the same dollar value as ACUs, which fixes the exchange rate between two units without pricing either. Anyone quoting a rate is either working from an order form we cannot see or from a historical rate card. If Cognition republishes rates, the arithmetic here should be redone with real values for M.
The 67% merge rate is vendor-reported and unaudited. It is Cognition describing its own product across an undisclosed population of customers and task types, published November 14, 2025. We use it because it is the best published figure and because being explicit beats being vague, not because we can verify it. The 15% figure from Answer.AI is a single team, twenty tasks, an early product version, and January 2025. Neither is a benchmark for your codebase. This is the single largest source of uncertainty in the model, and it is why the article's recommendation is to measure your own.
We could not verify US wage data from the primary government source. The Bureau of Labor Statistics Occupational Outlook Handbook returned 403 to every fetch method available to us on July 31, 2026, so the hourly figure here rests on the Stack Overflow 2025 Developer Survey — a large but self-selected sample — rather than on BLS. Treat $110/hour as a placeholder that your finance team should overwrite, not as an authority.
Every cost figure in this article is a model output, not an observation. We have not run Devin, have not reviewed its pull requests, and have not measured a merge rate or a review hour. Where a number appears without a source link, it is arithmetic on the assumptions stated in the model section.
The model omits real costs on both sides. It excludes onboarding time, environment configuration, the cost of a wrong change that merges, and the value of engineer attention freed for other work. It also excludes the possibility that review gets structurally cheaper as tooling improves — Devin Review's Auto-Fix is aimed squarely at that, and if it works as documented, R falls and the break-even table gets more forgiving still.
Prices move. Every figure was fetched on July 31, 2026 and Cognition changed its plan structure materially at some point before that date. Re-fetch before you build a business case.
Frequently asked questions
Devin's published plans, fetched from devin.ai/pricing on July 31, 2026, are Free at $0, Pro at $20 per month, Max at $200 per month, and Teams at $80 per month plus $40 per month per full dev seat. Enterprise is quoted and billed in Agent Compute Units at a rate set in the customer's order form. Pro and Max are single-user plans that cannot invite additional members, so any team of two or more starts at the $80 Teams minimum.
An ACU is an Agent Compute Unit, Devin's measure of work performed rather than time elapsed. Per Devin's usage documentation, consumption accrues on the number and complexity of actions the agent takes (planning, context gathering, task execution, browser actions, code execution) plus virtual machine time and networking bandwidth, which the docs describe as typically a small fraction of the total. Enterprise customers are billed in ACUs directly; self-serve customers consume a plan quota first and then prepaid on-demand credits.
Not on any page we could fetch. On July 31, 2026, devin.ai/pricing and docs.devin.ai publish no dollars-per-ACU figure for self-serve, no quota size, and no numeric overage rate — overage is described only as being "consumed at API pricing." Third-party articles widely repeat a Core-at-$20-plus-$2.25-per-ACU and Team-at-$500-with-250-ACUs ladder. Cognition's own documentation now files that structure under "Migrating from legacy ACU-based plans," noting that legacy Core plan users have been migrated to the Free plan and that on-demand credits carry the same dollar value as the ACUs they replaced. So the mechanism survives under a new name, but the rate does not appear anywhere public. Treat per-ACU figures from third parties as historical unless they cite a dated Cognition page.
Cognition does not publish a per-task average, but Session Insights publishes size bands that imply the range: XS is 2 ACUs or fewer, S is 5 or fewer, M is 10 or fewer, L is 20 or fewer, and XL is above 20. Enterprise thresholds are scaled by a factor of ten. Sessions landing at L or XL are flagged as unhealthy, meaning the scope was probably too broad. A well-scoped task that Cognition would call a good fit, meaning under three hours of human work with verifiable outcomes, should land in the XS to M range.
A full seat costs $40 per month, includes a Pro-equivalent daily and weekly usage quota, and grants Devin Desktop access. A flex seat is free, carries no fixed monthly charge, excludes Devin Desktop, and draws entirely from the team's shared pool of on-demand credits. Teams can have unlimited flex seats. Devin's own guidance is to give full seats to regular users and flex seats to occasional ones — which for most teams means the people who mainly review Devin's output rather than start sessions.
Yes, $80 is a genuine floor rather than a starting point that hides fees. Devin's self-serve documentation publishes the combinations: zero full seats means $80 charged as on-demand credits; one full seat means $40 of seat plus $40 of credits; two full seats means $80 with no credits included; and beyond two it is simply $40 per full seat. Below two full seats the remainder of the minimum arrives as prepaid credits the whole team can spend, and those credits roll over month to month, with purchased credits not expiring.
Enterprise admins can set per-Organization ACU limits from Enterprise Settings, and all Devin activity in that organisation stops when the limit is reached. Below Enterprise, the published controls are default session spending limits and auto-reload thresholds under Settings > Usage, plus a shared on-demand credit balance that any team member can draw from. Note that every one of these is denominated in ACUs, credits or sessions rather than dollars, and that none of them meters the human review time that this article argues is the larger cost.
That is the wrong comparison, and the right one is per task rather than per person. Cognition positions Devin at tasks that "would take a junior engineer 4-8 hrs of work." Modelling that as $440 to $880 of hand-built work at a $110 loaded hourly rate, and modelling agent cost as metered spend plus review hours divided by merge rate, break-even lands near a 9% merge rate at 30 minutes of review per attempt and near 26% at 90 minutes. Both are well below the 67% merge rate Cognition reports for itself. The risk in the deal is not the price; it is review time and reviewer availability.
Nobody can tell you honestly, and that is the point of measuring. Cognition reported in November 2025 that 67% of Devin's pull requests are merged, up from 34% the prior year — a vendor-reported figure across its own customers. Answer.AI, evaluating an early version in January 2025, reported 3 successes out of 20 attempted tasks. Your own rate will depend on test coverage, how clearly tasks are specified, whether your codebase has patterns to imitate, and how large the diffs get. A two-week trial on the $20 Pro plan will tell you more than any published figure.
No. Devin's usage documentation states that the agent does not consume usage while sleeping, while waiting for your response, while waiting for a test suite to run, or while setting up and cloning repositories, aside from the small amount required to keep the VM running. Sessions sleep automatically after roughly 0.1 ACUs of inactivity. The behaviour that actually inflates consumption is long, poorly scoped sessions with many corrective messages — which is also the behaviour that produces large diffs and low merge rates.
Buy Teams until you need the hard ceiling, then buy Enterprise for the ceiling rather than for the rate. The structural difference that matters is per-Organization ACU limits, which stop all activity when hit, along with SAML/OIDC SSO, centralised admin controls, VPC deployment and consumption analytics broken down by organisation and user. Because Cognition publishes no Enterprise rate card, no one can tell you from public data whether a negotiated ACU rate beats $40 per full seat plus credits — so go into that conversation with your own measured consumption and your own measured merge rate.
Ready to Govern Your AI?
Talk to LeapForce — one controlled layer for every AI tool, connector, model, and agent.
Comments