How much are our engineers actually spending on AI tokens each month?

Monthly token spend is your provider’s per-million-token rate multiplied by the tokens your team consumed, and the second number lives in usage exports most engineering orgs have never pulled. Nobody can quote you the figure, including us. The rate is published. The consumption is not, and it varies by a wide margin between two teams of the same size running the same tools.
Why the number is hard to find in the first place
The seat license arrives as one line on one invoice, once a month, and it does not move. Token spend arrives across three or four accounts, some of them opened by a staff engineer with a company card during a spike in June, and it moves every week.
An engineering director we would describe as well organised sat down last quarter to answer a board question about AI cost and found four places the money was leaving: the assistant subscription in the IT budget, a model provider account under engineering, a second provider account under the data team, and a per-request charge inside a code review integration that nobody had connected to token consumption at all.
None of that is negligence. The invoices are structured for the vendor’s billing model rather than for your budget lines.
How to work out the monthly number from your own invoices
Six steps, and the arithmetic is small once the inputs exist.
1. List every account that can consume tokens. Model provider consoles, the coding assistant, the code review bot, any internal service calling a model, and any cloud marketplace subscription. Ask finance for every vendor with the word AI in the payee name, then ask engineering what is missing from that list.
2. Pull the usage export, not the summary. Every major provider exposes token counts by day and by API key. The summary gives you dollars. The export gives you dollars and the reason for them.
3. Split the tokens you send from the tokens the model returns. Providers price the two directions differently, often by a factor of three or more. A workload that reads a large codebase and writes a short patch has a cost profile nothing like a chat session, and a single blended rate hides that completely.
4. Multiply by the published per-million rate for each model you use. Take the rate from your provider’s current pricing page rather than from a rate you remember, because the same model name at two versions can bill differently.
5. Divide by engineers with access. Not headcount. Access. This gives you a monthly cost per engineer with access, which is the figure a CFO will ask for by name.
6. Divide again by engineers who used it that month. The gap between step 5 and step 6 is the first thing worth acting on, and it is usually large.
We publish no benchmark for any of these figures. Any monthly dollar total in a blog post, ours included, would be a number invented to look authoritative next to your invoice.
Why the per-seat intuition breaks for tokens
Seat pricing has a ceiling you can calculate in advance. Multiply the rate by the headcount and you have next year’s worst case.
Token pricing has no such ceiling, because consumption is set by how the tool is used rather than by how many people hold a login. One engineer running an agent that reads 40 files, plans, edits, runs tests, and retries on failure can consume more tokens in a single afternoon than a colleague doing chat-style completions consumes in a month. Both engineers show up identically on a seat report.
That is the whole reason the monthly question keeps coming back. The cost driver moved from a count of people to a pattern of work, and most cost reporting still counts people.
Which usage patterns move the number most
Five patterns explain most of the variance between teams, and all five are measurable from the usage export once you know to look.
Context size per request. An agent given the whole repository as context pays for the whole repository on every turn. The same task scoped to the three relevant directories costs a fraction of it and often produces a better patch.
Turn count per task. Agentic workflows bill per step. A task that converges in 4 steps and a task that loops 30 times before hitting a step limit differ by an order of magnitude on the same invoice line.
Model choice per task. Teams frequently route every request to the largest available model, including lint fixes and commit message drafting.
Retry behaviour on failure. A failing test suite inside an agent loop can trigger repeated full-context retries. This is the single most common cause of a spike nobody can explain.
Batch versus interactive. Providers commonly discount asynchronous batch processing. Work with no human waiting on it, such as test backfills or documentation passes, is often billed at a lower rate if it is submitted that way.
If you want the spend question answered alongside what the team shipped in the same window, connect a repository and the historical window scores in the first pass.
What the spend buys, and how to read it
A monthly cost figure on its own invites exactly one response from finance, which is to make it smaller. The useful version of the number sits next to a measurement of what changed in delivery while that money was being spent.
The measurement to pair it with is the value of what shipped rather than the count of what shipped. Navigara’s Q2 2026 study, covering 699 engineers and 137,592 qualifying commits across 65 public repositories at 6 organisations, splits the annual change into its two components, and the split is the part worth copying into your own reporting.
Components of the Q2 2026 change in performance per engineer
| Component | Year over year | Quarter over quarter |
|---|---|---|
| Commits per engineer | +14.9% | -11.2% |
| Performance per commit | +100.4% | +32.4% |
Year over year and quarter over quarter, 699 engineers, 65 public repositories
Grouped columns comparing two components of the Q2 2026 result, with commits per engineer up 14.9 percent year over year against performance per commit up 100.4 percent, showing the annual change concentrated in the value of each change rather than the number of changes.
Commits per engineer moved 14.9% year over year. Performance per commit moved 100.4%. If you report token spend against commit volume, you are measuring the component that barely moved. The study also states plainly that it does not quantify what share of that level shift, if any, came from assistant adoption, and 21 of the 65 repositories in the sample are AI or agent SDKs whose category demand grew across the same window.
What to bring to the budget meeting
Four figures, each with a window and a source attached.
Total token spend for the month, summed across every account. Cost per engineer with access. Cost per engineer who used it. And the throughput measurement for the same period, with its confidence interval stated rather than implied.
The last one is what turns a cost conversation into a renewal conversation. A director who can say what the team delivered while the money was spent is answering a different question from a director who can only say how much the money was.
Talk to us if you want the spend and delivery sides read against the same window.
Frequently asked questions
- What is a normal monthly AI token spend per developer?
- There is no published figure we would stand behind, and we do not publish one. Consumption varies by more than an order of magnitude between teams of similar size depending on context size, agent step counts and model routing, so a median would mislead more than it helped. Pull your own usage export and calculate it from your provider’s current per-million-token rate.
- Why did our token bill increase without adding engineers?
- Almost always a change in how the tools are used rather than who uses them. Agentic workflows bill per step and per token of context, so switching a team from chat-style completions to an agent that reads a repository before editing raises consumption sharply at constant headcount.
- Does token spend show up as a per-developer figure?
- It can be divided by headcount, and that division answers a budgeting question. It does not answer a performance question, and a per-person token ranking rewards whoever burns the most context. Keep the cost view at team level, the same as the throughput view.
- How often should we review AI token spend?
- Monthly against the invoice, and quarterly against delivery. The monthly review catches runaway loops and forgotten accounts. The quarterly review is where the renewal decision gets made, and it needs a throughput baseline to be worth holding.
- Can we cap token spend without slowing the team down?
- Yes, mostly through scoping and routing rather than hard budget limits. Context scoping, model routing by task type, agent step caps and batch submission for asynchronous work reduce consumption without changing what an engineer can attempt.

