Blog

How do I calculate the ROI of our AI coding tools?

Jirka Bachel7 min read
Calculator and notebook beside a laptop.
Calculator and notebook beside a laptop. Photo by Mikhail Nilov · Pexels License

ROI on an AI coding tool is the throughput your team gained since adopting it, measured against your own pre-AI baseline, divided by everything the tool costs. Four steps: set the baseline, measure throughput after, price the full cost, divide. The arithmetic takes a minute. The baseline takes a quarter, and most teams skip it. Then they divide a number they like by a number they trust and call it a return.

What counts as ROI on an AI coding tool

Most engineering orgs running an AI ROI number today are running a satisfaction survey with a dollar sign in front of it. Developers say they feel faster. Cycle time looks shorter. Somebody multiplies a percentage by a loaded salary cost and the deck is done.

I’ve sat in that chair. As a former CTO, I’ve signed off on productivity numbers I couldn’t have defended under one follow-up question, because I had no measurement of what the team’s throughput looked like before the tools arrived.

ROI needs three quantities, and only one of them is easy:

  1. Throughput now. What the team ships, valued by the size and risk of the change rather than the count of commits.
  2. Throughput before. The same measurement, run against the repository history from before anyone installed an assistant.
  3. Total cost. Licenses, tokens, the infrastructure the tools call, and the review time the tools created.

The gap between 1 and 2, priced against 3, is the return. Everything else is a vibe with a spreadsheet attached.

How to set a pre-AI baseline when you did not plan for one

Good news on this one. Your baseline already exists and you’ve been storing it for years.

Your Git history is a complete record of what the team shipped, when, and at what size. Pick a window that ended before the first assistant landed in the workflow. For most US teams that means somewhere between late 2022 and mid 2023, depending on when Copilot moved from curiosity to default. Two full quarters is enough. One quarter will carry too much noise from a single large release.

Three rules for choosing the window:

  • Exclude a reorg. If the team changed size by more than about 20% inside the window, the baseline is measuring the reorg.
  • Exclude a migration. A monolith-to-services quarter produces a volume of change that has nothing to do with feature throughput.
  • Keep the same repositories. A baseline drawn from four repos and compared against nine repos isn’t a comparison.

Then measure the historical window with exactly the same method you’ll use going forward. This is the step that kills most internal attempts. Teams measure the past with one definition, the present with another, and the delta they report is mostly the definition change.

What to measure after the tools land

Commit count goes up when an assistant is in the loop. So do lines changed, pull request count, and any metric that counts events. None of that tells you whether more value hit production.

Measure the value of the change instead. Navigara does this with Engineering Throughput Value, which scores each merged change by what it actually was: a complex refactor, a performance fix, a feature unblock, a dependency bump. Three lines of code can be any of those. A metric that counts the three lines can’t tell them apart.

Two constraints matter here, and they’re the reason this doesn’t turn into a monitoring project:

Measure at team level. A per-developer throughput number is a leaderboard waiting to happen, and a leaderboard changes the behaviour it’s measuring within about two sprints. The unit that owns a delivery commitment is the team, so that’s the unit worth measuring.

Measure what shipped. Not keystrokes, not hours in the editor, not assistant acceptance rate. Acceptance rate tells you developers pressed tab. It doesn’t tell you the suggestion survived review.

If you want the baseline calculation run against your own repositories before you build the internal version, connect a repo and the historical window is scored in the first pass.

What an AI coding tool actually costs

The license line is the small one. The four costs that move an ROI number:

Seats. Whatever you pay per developer per month, times the seats you actually provisioned rather than the seats in use. Finance is paying for both.

Tokens. This is the line that surprises people, because it doesn’t behave like a seat. An agentic workflow that reads a large codebase before it writes anything can consume more tokens in an afternoon than a chat-style assistant does in a month. Token spend scales with how the tool is used, not with how many people use it.

Compute the tools trigger. CI runs on more pull requests. Test suites execute more often. Preview environments spin up more frequently. That bill lands in a different budget line and usually doesn’t get attributed back.

Review time. Assistants shift the bottleneck rather than removing it. More changes arrive at review, and senior engineers absorb the difference. Review debt is a real cost with a real salary number behind it, and it’s the one most ROI calculations leave out entirely because it never appears on an invoice.

Add those four. That’s the denominator.

How to do the division and report a confidence interval

The division itself is one line of arithmetic. Throughput delta, expressed in engineer-equivalents, times your loaded cost per engineer, divided by total tool cost.

The part almost nobody does is state the uncertainty, and it’s the part that makes the number survive a CFO.

Navigara’s own Q2 2026 study shows why. Year over year, performance per engineer came in at +130% across the open cohort of 699 engineers. Quarter over quarter, the same cohort came in at +17.5%, with a 95% confidence interval running from -0.7% to +38.8%. The report states the consequence plainly: “The quarter-over-quarter change cannot be distinguished from zero.”

The study also splits the result into its components, and this is the part worth stealing for your own reporting. Commits per engineer moved +14.9% year over year. Performance per commit moved +100.4%. Almost all of the gain sits in what each change was worth, and almost none of it sits in how many changes there were.

Where the Q2 2026 gain came from

ComponentYear over yearQuarter over quarter
Commits per engineer+14.9%-11.2%
Performance per commit+100.4%+32.4%

Year over year, 699 engineers, 65 public repositories

Grouped columns comparing two components of Navigara’s Q2 2026 result, showing commits per engineer up 14.9 percent year over year while performance per commit rose 100.4 percent, so the gain sits in the value of each change rather than in the number of changes.

Same dataset, same quarter, two windows, and one of them supports a claim while the other doesn’t. If you report the annual figure without the interval, you’ll be asked for the quarterly figure in the next meeting and you’ll have to walk it back. The study carries one more caveat worth copying into your own reporting: it doesn’t quantify what share of the change, if any, came from assistant adoption. Correlation is what the measurement gives you. Say so before someone else says it for you.

Report both. A number with a stated interval reads as instrumentation. A number without one reads as marketing, and your CFO has met marketing.

What a real ROI number looks like

It’s smaller than the vendor slide and more defensible than anything else in the room. It has a window on it, a repository list, a stated method, and an interval. It distinguishes the quarter where the gain was real from the quarter where the gain was inside the noise.

And it holds up when the question changes. “Is AI making us faster” is the question you get asked this quarter. “Should we renew at 3x the token spend” is the question you get asked next quarter, and only the baseline version of the number can answer it.

Why your first number will be wrong

It will be, and that’s fine. The first pass surfaces the problems in the data rather than the answer: a repository nobody remembered, a bot account inflating the historical window, a three-week gap where the CI provider changed and the commit metadata changed with it.

Fix those, rerun, and the second number is the one you take into the meeting. Talk to us if you want the first pass run against your history rather than building the pipeline yourself.

Frequently asked questions

How long does it take to calculate AI coding tool ROI?
The measurement itself takes a day once repositories are connected, because the historical data is already in Git. Getting a clean baseline takes longer, usually two to four weeks of resolving bot accounts, repository scope, and periods where the team changed size.
Can I calculate AI ROI without a pre-AI baseline?
You can produce a number. It compares your team to an industry average, which tells you little when your stack, team composition and constraints differ from the average. Your own history is the only comparison that controls for those.
Does measuring AI ROI mean tracking individual developers?
No. Throughput measurement works at team and repository level, and that’s the level at which delivery commitments are owned anyway. Per-developer reporting turns the measurement into a leaderboard and changes the behaviour it’s trying to observe.
What should I include in the cost side of the calculation?
Seat licenses at provisioned count, token spend, the CI and preview-environment compute the tools trigger, and the review time senior engineers now absorb. The last two are usually missing and they’re usually material.
Is cycle time a good proxy for AI ROI?
Cycle time gets shorter when changes get smaller, which is exactly what assistants encourage. Falling cycle time alongside flat throughput means the work got chopped up, not sped up. Read the two together or the trend misleads you.
What ROI figure should I expect?
Navigara doesn’t publish a customer ROI benchmark, so any figure here would be invented. The published research measures change in performance per engineer across 65 public repositories at six named companies, which is a different claim from a customer return, and 21 of those 65 repositories are AI or agent SDKs whose category demand grew over the same window. Run the calculation on your own history and report what it says.

More from the blog