Why our AI bill exploded when we moved off per-seat pricing

Because the cost driver shifted from headcount to a work pattern, and the work pattern changed in the same quarter. Seat pricing gave you a ceiling you could multiply out a year in advance. Usage pricing gives you a bill based on context size, agent step counts, and retries, none of which anyone was tracking when the contract was signed.
What per-seat pricing was quietly doing for you
Seat pricing was a forecasting instrument disguised as a line item. Headcount times rate times 12, and the number you got was the worst case. You could take it to a board meeting and defend it without knowing anything about how anyone used the tool.
That’s the part teams miss when they move. The seat model wasn’t cheaper. It capped the variance, and the cap was doing more work in your planning cycle than the rate was.
I’ve sat on the buying side of this. As a former CTO, I signed multi-year tool agreements where the entire finance conversation centered on the per-head rate, because it was the only variable. Nobody asked me how the tool would be used, and it wouldn’t have changed the number if they had.
What changed in the workflow at the same moment
Two shifts landed close enough together that most orgs experienced them as one event.
The first is agentic execution. A chat-style assistant answers a question and stops. An agent reads files, plans, edits, runs the test suite, reads the failure, and tries again. Each of those turns is a billable request carrying its own context.
The second is context size. Large context windows made it possible to hand a model a whole repository, and handing a model a whole repository became the default because it’s the easiest thing to do. You pay for that context on every turn, not once per task.
Put those together, and one engineer working the way engineers now work generates a request volume unrelated to the seat they occupy. Same person, same desk, 50 times the consumption of the colleague next to them who’s still using tab completion.
Where the money actually goes
Four drivers, and, in my experience, they explain almost every surprise invoice.
Context per turn. The most expensive habit in the building is giving an agent the entire codebase for a task that touches three files. It costs more and it frequently produces a worse patch, because the relevant context is buried.
Turn count. A task that converges in 4 turns and a task that loops 30 times before hitting a step cap differ by roughly an order of magnitude on the invoice. If you haven’t set a step cap, you don’t have an upper bound.
Retry storms. An agent loop wrapped around a flaky test suite will retry with full context each time. This is the single most common cause of a spike that nobody can account for at month-end, and it’s usually one repository doing it.
Model routing. Teams point everything at the largest model available, including commit message drafting and lint fixes. The rate difference between tiers is published on every provider’s pricing page, and almost no one compares it to their own task mix.
I’m not going to put dollar figures on any of that. I don’t have your provider’s contract, your negotiated rate, or your task mix, and a number I invented here would be worse than no number because you’d anchor on it. Take your provider’s published per-million-token rate, pull your usage export, and multiply it by your own consumption.
Why the invoice arrives before anybody can explain it
The billing cycle and the engineering cycle run at different speeds, and that gap is where the panic lives.
Picture the meeting. A finance partner has the September invoice open on a laptop, three times the August figure, and she’s asking a reasonable question: what did we get for it? The staff engineer in the room knows exactly which two weeks caused it, because he was the one who wired an agent into the migration work, and it saved the team a fortnight. He can say that. He can’t show it.
Nobody’s behaving badly in that room. The cost side has an instrument with daily granularity and the delivery side has an anecdote.
If you want the delivery side of that conversation measured over the same window as the invoice, connect a repository and the historical window scores in the first pass.
How to get an upper bound back
You can’t return to seat pricing, and you probably shouldn’t want to. You can rebuild the ceiling with four controls.
Step caps on every agent. A hard maximum on turns per task. This alone converts an unbounded workflow into a bounded one and exposes the tasks that really need more room, rather than hiding them inside a runaway loop.
Scoped context by default. Directory-level or package-level scope, with whole-repository context as an explicit choice somebody makes rather than the setting nobody changed.
Routing rules by task class. Cheap model for formatting, renaming, commit messages, and test scaffolding. Expensive model for design, debugging and anything touching a public interface.
Per-project budgets with account-level alerts. Not to block work. To make the spike visible on the day it happens rather than on the invoice five weeks later.
Those four give you a forecastable range. A range with stated assumptions is what finance wanted from the seat model in the first place.
What to measure on the other side of the bill
A cost number, on its own, has only one available response: to cut it. To hold a renewal, you need the delivery figure for the same window, and it has to be the shipped value rather than the count.
Our own Q2 2026 study splits the change in performance per engineer into exactly two components, across 699 engineers and 137,592 qualifying commits in 65 public repositories across 6 organizations.
Components of the Q2 2026 change in performance per engineer
| Component | Year over year | Quarter over quarter |
|---|---|---|
| Commits per engineer | +14.9% | -11.2% |
| Performance per commit | +100.4% | +32.4% |
Year over year and quarter over quarter, 699 engineers, 65 public repositories
Grouped columns show commits per engineer up 14.9 percent year over year, while performance per commit rose 100.4 percent, with quarterly figures at minus 11.2 and plus 32.4 percent, so the annual change reflects the value of each change rather than the count of changes.
Commits per engineer moved 14.9% year over year. Performance per commit moved 100.4%. Report token spend against commit volume, and you’re dividing a variable cost by the component that barely moved.
Two caveats travel with that, and I’d carry both into any meeting where you cite it. The study doesn’t quantify what share of the level shift, if any, was attributable to assistant adoption. And 21 of the 65 repositories are AI or agent SDKs whose category demand grew across the same window, so for those repositories, throughput per engineer and market growth can’t be separated.
What I’d do in the first week
Find every account that can consume tokens, including the one opened on a personal card during a spike. Set step caps. Turn on account-level alerts. Then run the delivery measurement backward over the same months the invoices cover, so the next time the bill moves, you have two lines on the same chart instead of one.
Talk to us if you’d rather have that second line before the next invoice lands.
Frequently asked questions
- Is usage-based AI pricing more expensive than per-seat?
- It depends entirely on how your team works, and that’s the point. A team using chat-style completions usually pays less under usage pricing. A team running agents over large context pays considerably more, at identical headcount.
- How do I forecast AI cost under usage pricing?
- Forecast a range rather than a figure. Take your provider’s published per-million-token rate, pull a 60- to 90-day usage export, and model a low case at current consumption and a high case at your worst observed week, extended across the quarter. State the assumptions next to the range.
- What caused our AI bill to triple in one month?
- Check three things in order: a new agentic workflow in one repository, a retry loop against a failing test suite, and whole-repository context on a task that didn’t need it. One of those three explains most of the sudden increases, and all three appear in the usage export by API key and day.
- Should we cap AI spend per developer?
- Cap the mechanics instead of the person. Step caps, context scope, and model routing reduce consumption without telling an engineer they’ve run out of budget mid-task, and they don’t create a per-person cost ranking that nobody can act on fairly.
- Does Navigara publish a benchmark for AI cost per engineer?
- No. We do not publish a cost-per-engineer benchmark, and any figure we quoted would be invented. The published research measures change in performance per engineer across public repositories at 6 named companies, which is a different claim from a cost benchmark.

