How to prove engineering productivity to a CFO who has never read a commit log

Give the CFO a number built the way every other number in that room is built: a quantity, a period, a comparison against your own prior result, and a stated error. A commit count has none of those properties and won’t survive the first follow-up question. A value-scored throughput figure with a named window, a named repository set, and an interval will be interrogated the same way revenue is.
What the CFO is asking for when they ask about engineering productivity
The question sounds like a challenge and it usually isn’t one. Engineering is the largest line in the model that finance can’t read directly. Sales has pipeline conversion. Marketing has cost per acquisition. Support has resolution time per ticket. Engineering arrives with a headcount, a cloud bill, a tooling bill that tripled in 18 months, and a board of tickets that means nothing to anyone who doesn’t attend the standup.
As a former CTO, I owned the answer to that question every quarter without an instrument that could produce it. The team was performing. I knew it from code review, from incident response, from the things that didn’t break. None of that converts into a row in a model.
A number finance trusts has four properties. A quantity. A period. A comparison. A stated error. Revenue has all four. Churn has all four. Sprint velocity has one, because the unit is invented by the team measuring it and recalibrated whenever the team feels slow.
Bring the four properties, and the meeting changes shape. You stop defending the department and start being asked what the number implies for next year.
Why a volume number fails in a finance conversation
Commits, pull requests, story points, and lines changed share one weakness. They count events. The count of events moves for reasons unrelated to what reached production.
Navigara’s Q2 2026 study separates the two. Across 699 qualifying engineers and 137,592 qualifying commits in 65 public repositories, commits per engineer rose 14.9% year over year. Performance per commit, which scores each merged change based on what the change was, rose by 100.4% over the same window.
Activity against value, Q2 2026 study
| Component | Year over year | Quarter over quarter |
|---|---|---|
| Commits per engineer | +14.9% | -11.2% |
| Performance per commit | +100.4% | +32.4% |
699 engineers, 137,592 commits, 65 public repositories
Grouped columns comparing activity with value in Navigara’s Q2 2026 study, where commits per engineer rose 14.9 percent year over year and fell 11.2 percent quarter over quarter, while performance per commit rose 100.4 percent and 32.4 percent over the same two windows, so almost all movement sits in the value of each change.
Read those two rows aloud in a budget meeting, and the implication is hard to miss. Almost all of the movement sits in what each change was worth. A volume metric can only see the 14.9%, so it would have to describe that year as roughly flat.
It fails in the other direction too, and that direction is the expensive one. Quarter over quarter in the same study, commits per engineer fell by 11.2%, while performance per commit rose by 32.4%. Anyone holding only the commit count would have reported a slowdown in that window. Carry the study’s own wording when you quote it: “The quarter-over-quarter change cannot be distinguished from zero.”
What to bring instead
Four items. Each one exists to survive a specific follow-up question.
A value-scored throughput figure, at team level. Each merged change scored by what it was: a complex refactor, a performance fix, a feature unblock, a dependency bump, a test backfill. 3 lines of code can be any of those. Navigara scores changes this way through Engineering Throughput Value, and the reporting unit is the team, because the team owns the delivery commitment.
Your own prior result as the comparison. Not an industry average. The follow-up question is always “compared to what”, and the only comparison that controls for your stack, your team composition, and your constraints is your own history. Git already holds it.
A named window and a named repository set. Written on the slide. The second follow-up question is “why that period,” and a period chosen after seeing the result is a period the CFO will discount to zero.
An interval. More on that below, because it’s the part that decides whether you get asked back.
If you want the baseline and the current window scored against your own repositories before you build an internal version, connect a repository and the historical window scores in the first pass.
How to price the engineering side in units finance already uses
The throughput figure is the numerator. Finance will ask for the denominator within about 90 seconds, so bring it assembled.
Loaded cost per engineer. Get it from finance rather than calculating it yourself. Using their figure removes an argument you gain nothing from winning.
Seats provisioned, not seats in use. Whatever the AI assistant costs per developer per month, multiplied by what was provisioned. Finance is already paying for the idle ones and will know the number.
Token spend. This line doesn’t behave like a seat. An agentic workflow that reads a large codebase before writing anything can consume more tokens in an afternoon than a chat-style assistant uses in a month. It scales with how the tool is used rather than with how many people hold a license.
Compute the tooling triggers. More pull requests means more CI runs, more test executions, more preview environments. That bill lands in a different budget line and usually never gets attributed back to the tool that caused it.
Review time. More changes arrive during review, and senior engineers absorb the differences. It has a salary number associated with it, but it appears on no invoices, which is why it’s missing from most internal ROI attempts.
One thing to leave out: a return figure borrowed from a vendor. Navigara publishes no customer ROI benchmark, so any such number would be invented, and a CFO who checks one unsourced figure will discount every other figure on the page.
How to state uncertainty without losing the room
Engineering managers tend to hide intervals, on the theory that a range looks weaker than a point. In a finance meeting, the opposite holds. Every number the CFO already trusts arrived with a range attached.
The Q2 2026 study is a usable model here. Year over year, performance per engineer was 130% across the open cohort of 699 engineers, and +82% across the fixed panel of 388 engineers active in all 6 quarters. Quarter over quarter, the open cohort increased by 17.5%, with a 95% confidence interval ranging from -0.7% to +38.8%. The report states the consequence without softening it: “The quarter-over-quarter change cannot be distinguished from zero.”
Two habits worth copying from that. State the interval next to the point estimate, and say which window supports a claim and which one doesn’t. The study also declines to attribute the change to any cause: “The study does not quantify what share of the level shift, if any, is attributable to AI coding assistant adoption.” Correlation is what a measurement of merged changes can give you. Say it first, because the alternative is hearing it from the CFO.
A number with an interval reads as instrumentation. A number without one reads as advocacy, and your CFO has been reading advocacy all week.
What to do when the CFO asks for the per-developer version
Somebody will ask, usually for reasonable motives. Productivity in every other department is reported per head, so per head is the shape finance expects.
Answer with the mechanics. Per-developer commit volume has existed in Git for 20 years, and anyone could have built the ranking on a Saturday afternoon. Engineering orgs don’t, because that ranking penalizes hard refactors, rewards renaming variables, and inverts under any serious review of what shipped. A metric that reverses the moment it becomes visible is a liability in a budget model, which is an argument finance understands better than an ethics argument.
Then give them what they were reaching for. A value-scored change history makes the quarter’s largest piece of work legible as such, so recognition becomes possible without ranking anyone.
The four sentences that carry the meeting
Write these before you go in. Fill the brackets from your own data.
“Across [named repositories], between [start] and [end], team throughput moved [figure] against our own prior-year window.”
“Almost all of that movement is in the value of each change rather than in the number of changes, which is why our commit count looks flat.”
“The quarterly comparison sits inside the noise, so I’m not making a claim on it, and the quarter-over-quarter change cannot be distinguished from zero.”
“Against [total tool and compute cost], that puts the return at [figure], and I can show the method for every part of it.”
Four sentences, no adjectives, one admission of what the data can’t support. The admission is the part that buys the rest.
Talk to us if you want the first pass run against your own history before the next budget cycle starts.
Frequently asked questions
- What engineering metrics do CFOs care about?
- A quantity with a period, a comparison, and a stated error. In practice, that means throughput measured against your own prior window, priced against loaded engineer cost plus tooling and compute spend. Team-level figures, since teams own delivery commitments and financial plans.
- Is velocity or story points enough for a finance conversation?
- No, because the unit is defined locally and recalibrated by the team that reports it. A CFO who learns that points were re-scoped mid-year will discount the entire series. Use a measure whose definition doesn’t change when the team’s mood does.
- How do I answer “can you prove engineering is productive”?
- Give the comparison against your own baseline, name the window and repositories, and state the interval. Then state clearly which claims the data supports and which it doesn’t, because unsupported claims are wasted first.
- Do I need a pre-AI baseline before talking to finance?
- It helps considerably, and it already exists in your Git history. Pick 2 full quarters that ended before the first assistant entered the workflow, exclude any window that contains a reorg or a migration, and measure that period using the same method you’ll use going forward.
- What ROI figure should I quote for our AI tools?
- Your own. Navigara publishes no customer ROI benchmark, and the published research measures change in performance per engineer across 65 public repositories at 6 named companies, which is a different claim from a customer return. 21 of those 65 repositories are AI or agent SDKs whose category demand grew across the same window.
- Does reporting productivity to finance mean monitoring developers?
- No. The reporting unit is the team and the repository, which is also the unit finance plans with. Per-developer throughput reporting produces a ranking that misreads shared work, so it stays out of the reporting entirely.

