Blog

What one developer’s AI usage costs for 50, 100, and 500 engineers

Navigara7 min read
Jensen Huang at SC18.
Jensen Huang at SC18. Photo by Raysonho @ Open Grid Scheduler / Scalable Grid Engine · CC0 1.0

The per-developer figure varies with the distribution of how people work, so it behaves differently across those three sizes even when the underlying rates are identical. Below is the model, not the answer. We publish no cost-per-developer benchmark, and a dollar total printed here would be invented, so the arithmetic is written for you to run against your own provider invoice.

Why one cost-per-developer figure stops working as the team grows

At every size, token consumption is unevenly distributed, and the shape of that distribution is what your average hides.

A small number of engineers run agentic workflows over large context and consume most of the tokens. A larger number use chat-style completions and consume very little. Some provisioned seats consume nothing at all. Divide the invoice by headcount, and you get a figure that describes none of those three groups; then plan next year on it.

The distribution also changes with headcount, because the mechanisms that set consumption change. At 50 engineers, the pattern is set by individuals. At 500 it is set by platform defaults, and a default is a multiplier.

The model to run on your own numbers

Five inputs, all of which you already have or can pull this week.

R, the rate. Your provider’s current published price per million tokens, taken separately for the tokens you send and the tokens the model returns, and separately for each model tier in use. Take it from the pricing page today rather than from memory.

X, tokens per task. From your usage export, total tokens divided by task count for a representative fortnight. Split this by workflow if you can, because an agentic task and a completion differ by orders of magnitude and a blended figure will mislead you at every team size.

T, tasks per active engineer per day. Count tasks, not requests. A 30-turn agent run is one task.

A, the active share. The fraction of engineers with access who used the tools in the month. This is the number most teams guess and most teams guess high.

D, working days. Whatever your calendar says for the month.

The monthly cost for one active engineer is X times T times D times R. The monthly cost per engineer with access is that figure times A. Total monthly spend is the per-access figure times the count of engineers with access.

Run it twice. Once with your median task, once with your heaviest observed workflow. The gap between those two results is your forecast range, and it is the honest thing to take to finance.

What changes at 50 engineers

Consumption is concentrated in named individuals, and you can list them.

Two or three engineers have wired agents into the work that touches the most files, and they account for a large share of the invoice and a large share of the hard changes. The active share, A, is usually high because a 50-person org adopts tools by word of mouth rather than through a rollout.

The practical consequence is forecast variance. One person changing how they work, or one engineer leaving, moves your total by a visible percentage. A budget built on last month’s average will be wrong in either direction, and the error will be attributed to the tool rather than to the model.

At this size, run the model per workflow rather than per engineer, and track the count of agentic workflows in production. That count is your cost driver.

What changes at 100 engineers

Two or three distinct usage populations appear, and the average now sits between them rather than describing any of them.

Teams have diverged. The platform team runs agents over large context because the work spans packages. A product team runs short completions because the work is contained. Consumption per engineer between the two groups can differ by a factor of 20, and both groups appear identical in the seat report.

The active share starts to fall here because access is granted by default at onboarding, while use remains voluntary. That gap between engineers with access and engineers who used it last month is the first reduction most orgs at this size find.

Run the model once per team with its own X and T, then sum. A single organization-wide average across 100 engineers produces a plan that overfunds the product teams and underfunds the platform team, a combination that generates a mid-quarter emergency.

What changes at 500 engineers

Defaults dominate individuals, and the largest cost decisions are no longer made by anyone in particular.

The default context scope in a shared agent configuration, the default model in a shared routing rule, the absence of a step cap in a template that 40 teams copied. Each of those multiplies across the whole population. At this size, a configuration change in a single internal template can move the monthly invoice more than a hiring round does.

Negotiated rates enter the picture too, which breaks any comparison against published pricing, including any benchmark. Committed-use discounts and enterprise agreements mean your R is not the public R.

Run the model per usage population rather than per team, with a named owner for each platform default. Then run a sensitivity pass on the defaults: what the invoice does if context scope doubles, if the step cap is removed, if routing sends 20% more traffic to the larger model. Those three numbers are the ones a CFO at this size should be shown.

If you want the delivery side measured over the same window as the cost model, connect a repository and the historical window scores in the first pass.

Which inputs move the answer most

Ranked by how much movement we would expect from a plausible change in each, tokens per task comes first, because context size and turn count both fall under it and can each change by an order of magnitude without anybody filing a ticket.

Active share comes second and is the cheapest to fix, since it is a provisioning issue rather than an engineering one.

Rate comes third, and it is the one teams spend the most time meeting about. Moving a tier of traffic from the largest model to a smaller one changes R for that traffic, which is real, and it is smaller than the change available from properly scoping the context.

Tasks per day comes last. It is roughly stable, per the engineer, and it is the input you least want to reduce, because reducing it means the team is attempting fewer things.

What the cost model needs next to it

A cost model alone answers one question: how much. Renewal decisions need the other side of it, measured over the same window and at team level.

Navigara’s Q2 2026 study measured changes in performance per engineer across 65 public repositories across 6 organizations, and it reports on two cohorts, which is directly relevant to anyone modeling across different team sizes.

Q2 2026 change in performance per engineer, open cohort against fixed panel

ComparisonOpen cohortFixed panel
Year over year+130%+82%
Quarter over quarter+17.5%+3.5%

Open cohort: 699 engineers; fixed panel: 388 engineers continuously active across six quarters

Grouped columns comparing two cohorts in the Q2 2026 study, where the open cohort of 699 engineers rose 130 percent year over year, compared with 82 percent for the fixed panel of 388, and the quarterly figures were 17.5 percent and 3.5 percent, respectively.

The open cohort came in at +130% year-over-year. The fixed panel of 388 engineers continuously active across all six quarters came in at +82%. Same study, same window, and the difference is cohort composition, which is the same effect that makes a per-developer cost average unreliable when the underlying population is changing.

The quarterly figure carries a 95% confidence interval from -0.7% to +38.8%, and the report states plainly: “The quarter-over-quarter change cannot be distinguished from zero.” The study also does not quantify what share of the level shift, if any, was attributable to assistant adoption, and 21 of the 65 repositories are AI or agent SDKs whose category demand grew over the same window.

Build the cost model on your invoice. Build the delivery figure on your own repository history over the same months. Then the renewal conversation has two lines on one chart and a stated interval on the second.

Talk to us if you want the delivery side run against your history before the next budget cycle.

Frequently asked questions

What does AI cost per developer per month?
We publish no benchmark, and any figure here would be invented. Calculate it as tokens per task × tasks per day × working days × your provider’s published per-million-token rate, then multiply by the share of engineers who used the tools that month.
Does the AI cost per developer decrease as the team grows?
Not reliably. Negotiated rates and committed-use discounts push the rate down, while shared-platform defaults on context scope and model routing push consumption up; at 500 engineers, the defaults usually move more money than the discount does.
Why is our average cost per developer so different from that of another company?
Because consumption is set by workflow rather than headcount, two orgs of the same size, running different context scopes and agent step limits, can differ by more than an order of magnitude. Negotiated rates also break the comparison, since neither company is paying the published price.
Should we budget per engineer or per team?
Per team at 100 engineers and above, per usage population at 500. A single organization-wide average overfunds teams doing contained work and underfunds teams doing cross-package work, which is the pattern that produces a mid-quarter budget escalation.
How do we know the spend is worth it at any size?
Measure the value of what shipped over the same window as the invoice, at team level, against your own pre-AI baseline rather than an industry average. A cost figure with a delivery figure next to it and a stated confidence interval survives a finance review, and a cost figure on its own invites a cut.

More from the blog