Blog

12 questions your CEO will ask about AI ROI, and how to answer them

Navigara7 min read
Dario Amodei at TechCrunch Disrupt 2023.
Dario Amodei at TechCrunch Disrupt 2023. Photo by Kimberly White / TechCrunch · CC BY 2.0

Answer with your own pre-AI baseline, a stated window, and a stated interval, and say plainly which questions your data cannot answer. Four of the twelve below fall into that second group. Naming them is what makes the other eight credible, because a CEO who catches one overreach discounts everything else in the deck for the rest of the year.

What all twelve questions have in common

A VP of engineering sat in a Tuesday leadership meeting last spring, where a slide showed a 40% productivity gain from the assistant rollout. The CEO asked where the number came from. It came from a developer survey. The meeting moved on, and the slide never appeared again.

The questions below are the ones that arrive when the number does hold. Each answer here is written to survive one follow-up, which is the only test that matters in that room.

Questions about the money

What did we get for the spend?

Give the throughput change against your own pre-AI baseline, stated as a percentage with the window attached, then the total cost including seats, tokens, the CI and preview-environment compute the tools trigger, and the senior review time the extra changes consume. The last two rarely appear on an invoice and they are usually material. One ratio, two dates, one repository list.

Can we cut headcount now?

The data does not support that conclusion, and saying so protects you twice. A throughput measurement tells you what the team shipped and the value of each change. It says nothing about what the team would ship at 80% of its current size, because capacity, review depth, and on-call all degrade non-linearly. If the CEO wants a headcount case, it needs a demand forecast, not a productivity ratio.

What happens if we stop paying for it?

Unknown until you measure it, and a small reversal is a real way to find out. Nothing in a before-and-after comparison establishes what happens on removal, because teams retain new working habits and the codebase has already changed shape. A 6- to 8-week pause on one team, measured the same way as everything else, gives a defensible answer. An estimate in the meeting does not.

Are we paying for seats nobody uses?

Some of them, yes, and this is the one question you can answer to the decimal today. Provisioned seats versus active seats is a billing query. Token spend is the line that surprises people, because an agentic workflow reading a large codebase can consume more in an afternoon than a chat-style assistant does in a month. Token cost scales with how the tool is used rather than with how many people hold licenses.

Questions about the number itself

Why is our number smaller than the vendor claimed?

Vendor figures usually measure task completion time in a controlled exercise. Your figure measures the number of merged changes reaching production through your review process, test suite, and release cadence. Those are different quantities, and the second one is the one your board is asking about. A smaller number with an attached method is the stronger position.

Can you prove the tools caused this?

No, and no measurement of your history can. Navigara’s own Q2 2026 study, covering 699 engineers and 137,592 qualifying commits across 65 public repositories at six organizations, carries a direct caveat: “The study does not quantify what share of the level shift, if any, is attributable to AI coding assistant adoption.” Correlation across a window during which many things changed is what this class of measurement gives. Say it before someone else says it for you.

Why did it drop this quarter?

Check whether the movement exceeds the noise in the series before explaining it. The Q2 2026 study is the illustration: year-over-year, performance per engineer came in at +130% across the open cohort, while quarter-over-quarter came in at +17.5%, with a 95% confidence interval of -0.7% to +38.8%. The report states the consequence: “The quarter-over-quarter change cannot be distinguished from zero.”

Same cohort, two windows, Q2 2026

WindowChange in performance per engineer
Year over year+130%
Quarter over quarter+17.5%

Open cohort of 699 engineers, 65 public repositories

Two columns comparing Navigara’s Q2 2026 results across two windows, with the year-over-year change at plus 130 percent and the quarter-over-quarter change at plus 17.5 percent, where the quarterly figure has a 95 percent confidence interval from minus 0.7 to plus 38.8 percent and therefore cannot be distinguished from zero.

A single quarter of deceleration in a noisy series is not yet evidence of a plateau. The same study says as much about its own data.

How much of this would have happened anyway?

Some of it, and the share is unknown. Team composition changes, a codebase that got easier to work in, a quarter with less incident load, and simple demand growth all move the same number. The Q2 2026 study runs into this in its own sample: 21 of the 65 repositories are AI or agent SDKs, unevenly spread: OpenAI 8 of 8, Microsoft 5 of 12, Vercel 3 of 9, Google 3 of 12, Cloudflare 2 of 14, and demand for that category grew over the same window, so for those repositories, throughput per engineer and market growth are not separable.

If you want your own baseline and current window scored the same way before the next board meeting, connect a repository and the historical window will be scored in the first pass.

Questions about everyone else

How do we compare to our competitors?

You cannot know that from this data, and neither can anyone selling you a comparison to it. Your competitors' private repositories are private. What can be compared is public-repository activity at named companies, which the Q2 2026 study covers for Cloudflare, Vercel, OpenAI, Google, Meta, and Microsoft, and peer-cohort benchmarking on Navigara’s Pro and Private Benchmark tiers, which references 200+ teams. Neither is your competitor’s internal delivery rate.

Are we behind on adoption?

Adoption rate is the wrong quantity to be anxious about, because it is easy to move and tells you little. Per-organization results in the Q2 2026 study grew year over year: Vercel +151.1%, Google +138.0%, Microsoft +123.2%, Cloudflare +96.4%, Meta +38.0%, and OpenAI +230%, all rebased from a Q3 2025 baseline. Public repositories only, with no view into private work or code review depth. Six organizations with broadly similar access to the same tools ended up six-fold apart, so the interesting variable lies elsewhere than whether the licenses were bought.

Questions about what happens next

What do you need from me?

Be ready with a short list, because this question closes. A decision on the repository set, two to four weeks to resolve the baseline, and one commitment: that the throughput figures go to finance and planning and stay out of individual performance reviews. That last one costs nothing and determines whether the team treats the measurement as an instrument or an audit.

When will this number be reliable?

Two to four weeks for a defensible baseline, and two to three quarters before quarterly movements mean much. The first pass surfaces the problems in the data rather than the answer: a repository nobody remembered, a bot account inflating the historical window, a gap where the CI provider changed and the commit metadata changed with it. The second number is the one that goes in the deck.

Why the four refusals are the valuable part

Four of these twelve answers decline to give a figure. Causation, competitor comparison, the counterfactual, and what happens on removal.

A CEO who hears “we cannot know that from this data, here is what we can know, and here is what it would take to find out” gets a clear signal about which numbers in the deck are load-bearing. A deck where every question gets a confident percentage gives no such signal, and the first one to break takes the rest with it.

The engineering manager in that room already knew what the team delivered. The work of the measurement makes it readable to people who were not in the room, with the uncertainty intact.

Talk to us if you want the baseline calculation run against your own repository history before the next board cycle.

Frequently asked questions

What is the best way to report the ROI of an AI coding tool to a CEO?
One ratio with a window, a repository list, a stated method, and a confidence interval, plus an explicit list of the questions the data cannot answer. The interval is what makes it read as instrumentation, and CEOs tend to trust a smaller number with stated uncertainty over a larger one without.
Can AI ROI measurement prove the tools caused the improvement?
No. Navigara’s Q2 2026 study states that it does not quantify what share of the level shift, if any, is attributable to AI coding assistant adoption, and the same limitation applies to any measurement of a single organization’s history. What you get is a correlation across a window, with its limits stated.
Should I benchmark our AI ROI against competitors?
Competitor internal delivery data is not available to anyone, so any such benchmark is built on public repositories or a peer cohort rather than on your competitors. Your own pre-AI baseline is the comparison that controls for your stack, team, and constraints.
How do I answer a headcount-reduction question using this data?
Say the data does not support a headcount conclusion and explain why: it measures what was delivered, not what would be delivered by a smaller team with less review capacity and the same on-call load. Then offer what would answer it: a demand forecast alongside the throughput figures.
What costs should be included in the AI ROI denominator?
Seat licenses at provisioned count rather than active count, token spend, the CI and preview-environment compute the tools trigger, and the senior review time that extra changes consume. The last two usually go unattributed and usually change the answer.
How long before quarterly AI ROI numbers are meaningful?
Expect two to three quarters. The Q2 2026 study shows why: its quarter-over-quarter change of +17.5% carried a 95% confidence interval from -0.7% to +38.8%, and the report states that the change cannot be distinguished from zero. Annual comparisons stabilize well before quarterly ones.

More from the blog