Blog

What to report to the board about AI’s impact on engineering

Navigara7 min read
Business leaders in a board meeting.
Business leaders in a board meeting. Photo by Werner Pfennig · Pexels License

Report four things: throughput measured against your own pre-AI baseline; the full cost of the tooling that produced it; how the mix of work changed; and one line stating what the data can’t prove. A board asks about AI in engineering because it’s deciding on next year’s tooling spend and headcount plan. Both decisions need a defensible figure with its uncertainty attached, and neither survives a number that gets walked back at the following meeting.

What the board is deciding when it asks about AI in engineering

The question arrives as curiosity and lands as a budget decision. Directors have read that AI coding assistants are changing engineering economics, and they want to know whether that has happened here.

There’s usually one person in the room driving it. A director who chairs the audit committee, with the tooling line item circled in the pack, asking why AI spend increased 4x while the delivery dates shifted by 2 weeks. That question deserves a real answer, and the answer will determine whether next year’s budget treats engineering as an investment with a measurable return or as a cost center with an unexplained increase.

A board can act on three kinds of statements. This changed by this much, compared to this. This cost this much. This is what we don’t know yet. Anything outside those three shapes gets politely absorbed and quietly discounted.

The four things worth reporting

Throughput against your own pre-AI baseline. Your Git history records what shipped before any assistant entered the workflow. Pick 2 full quarters that ended before adoption, keep the same repository set on both sides, and measure both windows with the same method. Teams that measure the past one way and the present another mostly report a change in definition.

Total cost, not the license line. Seats at provisioned count, token spend, the CI and preview-environment compute the tooling triggers, and the review time senior engineers now absorb. Token spend is the line that surprises boards, because it scales with how agentic the workflow is rather than with how many people hold a license.

The work mix shift. Where capacity went, split by features, tests, fixes, docs, and maintenance. This is frequently the most useful slide in the pack, because a board can read a shift in the mix even when the total is ambiguous.

One line on what the data can’t establish. Named explicitly, in writing, on the slide. It’s the line that makes the other three credible.

What the slide says

One slide, copyable. Fill the brackets from your own data and keep the structure.

Title: AI in engineering, [quarter] against [pre-AI baseline window]

Line 1, the change. “Team throughput per engineer moved [figure] against our [named window] baseline, across [named repository set].”

Line 2, the decomposition. “Commit volume per engineer moved [small figure]. The value per merged change moved [larger figure]. The change is in what each piece of work was worth rather than in how much work there was.”

Line 3, the cost. “Tooling and compute attributable to AI workflows: $[figure] for the period, of which $[figure] is token spend and $[figure] is added CI and preview compute.”

Line 4, the interval. “Annual comparison: [figure], 95% CI [lower] to [upper]. Quarterly comparison: [figure], 95% CI [lower] to [upper].”

Line 5, the caveat, verbatim on the slide. “The quarter-over-quarter change cannot be distinguished from zero.”

Line 6, attribution. “We can’t yet separate the effect of assistant adoption from hiring, scope change, and market demand. The figure is a measurement, not a causal claim.”

Line 7, the ask. One sentence naming the decision you want from the board, whether that’s renewal at a higher token budget, a headcount number, or 2 more quarters of measurement before either.

Seven lines, no adjectives. A director who reads that slide can ask a hard question and get a real answer, which is the entire point of putting a number in a board pack.

Why a volume figure doesn’t belong on the slide

Navigara’s Q2 2026 study shows what happens when activity and value get treated as the same quantity. Across 699 qualifying engineers and 137,592 qualifying commits in 65 public repositories across 6 organizations, commits per engineer rose 14.9% year over year, while performance per commit rose 100.4%.

Activity against value, Q2 2026 study

ComponentYear over yearQuarter over quarter
Commits per engineer+14.9%-11.2%
Performance per commit+100.4%+32.4%

699 engineers, 137,592 commits, 65 public repositories, 6 organizations

Grouped columns show that in Navigara’s Q2 2026 study, commits per engineer rose 14.9 percent year over year and fell 11.2 percent quarter over quarter, while performance per commit rose 100.4 percent and 32.4 percent over the same periods, so activity counts and delivered value move independently.

A board pack built on commit volume would have described that year as broadly unchanged. The quarterly column is worse: volume fell 11.2% while value per change rose 32.4%, so a volume slide reports a decline in a window where each change got more valuable. Either reading is wrong, and both are wrong in ways that only become visible after the board has already made a decision on it.

The decomposition is the defensible move. Put both components on the slide, and the board can see which one carries the result.

How to report uncertainty to a board without sounding uncertain

Intervals get left off board slides because a range looks weaker than a number. Directors read intervals all day in every other part of the pack, and a point estimate with no range reads as unfinished.

The Q2 2026 study is a usable model. Year over year, performance per engineer was +130% across the open cohort of 699 engineers and +82% across the fixed panel of 388 engineers who were continuously active in all 6 quarters. Two cohort definitions, two answers, both reported. Quarter over quarter, the open cohort came in at +17.5%, with a 95% confidence interval from -0.7% to +38.8%, and the report states the consequence without softening it: “The quarter-over-quarter change cannot be distinguished from zero.”

Three habits to copy.

Report the sample definition next to the figure. A cohort that changes composition between windows produces a different number from a fixed panel, and the gap between +130% and +82% in the same study is the size of that effect.

Separate the window that supports a claim from the window that doesn’t. The annual comparison in that study carries a claim. The quarterly one carries a direction and nothing more.

Say what the measurement can’t attribute. The same report states: “The study does not quantify what share of the level shift, if any, is attributable to AI coding assistant adoption.” A board that hears your limitations stops looking for the one you hid.

What to leave off

Per-developer figures. Never in a board pack. A ranking of engineers in a document that circulates to directors creates governance and retention problems, and misreads shared work as well. One engineer unblocks 4 others, one review prevents 3 incidents, and pair work produces a single commit with 2 contributors. Report at the team and repository levels, which is also the level at which the board plans headcount.

Assistant acceptance rate. It tells you developers pressed tab. It doesn’t tell you whether the suggestion survived review, and a board that learns the distinction later will apply the discount retroactively to everything else.

Vendor benchmarks presented as your result. Navigara publishes no customer ROI figures, and any number lifted from a vendor slide refers to somebody else’s repositories. One unsourced figure in a pack costs you the rest.

A single quarter treated as a trend. A quarter of deceleration in a noisy series isn’t yet evidence of a plateau, and a quarter of acceleration isn’t yet evidence of a return.

If you want the baseline window and the current quarter scored against your own repositories before the pack goes out, connect a repository and the historical window scores in the first pass.

What to say when a director asks for the headcount implication

The question follows the slide almost every time, and no published research answers it. Navigara doesn’t publish a headcount outcome, and the Q2 2026 study measures changes in public repositories resulting from merges rather than staffing decisions.

What the data does support is a statement about the work. Here’s what the team delivered, here’s what it cost including tooling, here’s what’s still queued behind current capacity. The board can make a staffing decision from that, which is their decision to make.

The one thing worth adding is the instrument’s boundary. Merged changes don’t capture the depth of code review, incident response, planning, or mentorship, and those activities absorb a large share of a senior engineer’s week. A headcount model built only on merged changes will underprice exactly the people the company would most regret losing.

Talk to us if you want the baseline and current-quarter read prepared against your own repositories before the next board cycle.

Frequently asked questions

What should an engineering AI update include in a board pack?
Throughput against your own pre-AI baseline; the full cost of the tooling, including token and compute spend; the shift in work mix across features, tests, fixes, docs, and maintenance; and an explicit line on what the data can’t attribute. One slide, with the window and repository set named.
How often should we report AI impact to the board?
Annually for the comparison that carries a claim, with quarterly updates reported as direction only. A quarter of movement in a noisy series isn’t yet evidence of a trend, in either direction, and reporting it as one costs credibility when it reverses.
Can we claim a specific ROI on AI coding tools to the board?
Only against your own baseline and your own cost lines. Navigara publishes no customer ROI benchmarks, so any figure from a vendor reflects different repositories, team compositions, and constraints. Report your own number with its interval instead.
Should individual developer metrics ever reach the board?
No. Board reporting belongs at the team and repository levels, where delivery commitments and headcount plans exist. Per-person figures in a circulating document misread shared work and create problems unrelated to measurement.
What do we say if the number is inside the noise?
Say that, and name what it would take to resolve it: more quarters, a fixed team comparison, or a wider repository set. Navigara’s own Q2 2026 study does this with its quarterly figure, reporting +17.5% with a 95% confidence interval from -0.7% to +38.8% and stating that the quarter-over-quarter change cannot be distinguished from zero.
How do we handle a board member who quotes a public AI productivity statistic?
Ask which sample it came from and over what window. Published figures vary widely with cohort definition, repository mix, and measurement method. In the Q2 2026 study, the same data yielded +130% for an open cohort of 699 engineers and +82% for a fixed panel of 388, showing how much the definition shifts the result.

More from the blog