Is our AI assistant actually making the team faster? Here is how to check

Check it with four readings taken together: throughput measured against your own pre-AI quarters; the split between how much activity changed and how much the value of each change changed; cycle time read alongside change size; and where the capacity went by type of work. Any one of those alone will mislead you. All four, at once, give you an answer you can defend when somebody pushes back.
What faster has to mean before you can check it
Picture the sprint review where somebody asks the question directly. A staff engineer, three months into the rollout, says the typing is faster, and the thinking is the same, and she isn’t sure which one the company bought.
That’s the measurement problem in one sentence. “Faster” can mean five different things, and the ones that are easy to measure matter least.
It can mean more changes merged per week. It can mean each change is worth more. It can mean that the gap between starting a ticket and shipping it has gotten shorter. It can mean the share of the quarter going to maintenance dropped. It can also mean developers experience less friction, which is real, worth knowing, and completely separate from delivery.
Pick which one you’re claiming before you open a dashboard. Most disagreements about AI productivity are two people measuring against different definitions.
Does throughput exceed your pre-AI quarters?
This is the first reading and the only one that answers the question as a CFO will ask it.
Take two quarters that ended before any assistant entered the workflow, score them with the same method you’ll use now, and compare. Same repository set, same contributor rules, bots excluded from both windows. The comparison is worth exactly as much as the method consistency behind it.
Industry averages can’t do this job. Your review culture, merge policy, and test burden are yours, and the average of six other companies still has none of them.
I spent years as a former CTO answering this question with conviction and no instrument, which works until the second follow-up question.
Did activity rise, or did the value of each change rise?
Second reading, and the one that changes how people think about the first.
Activity counts go up when an assistant is in the loop, almost mechanically. More commits, more lines touched, more pull requests. That movement is easy to find and easy to over-read.
Navigara’s Q2 2026 study split the change into its two components across 699 engineers and 137,592 qualifying commits.
The two components of the Q2 2026 change
| Component | Year over year | Quarter over quarter |
|---|---|---|
| Commits per engineer | +14.9% | -11.2% |
| Performance per commit | +100.4% | +32.4% |
Open cohort, 699 engineers, 65 public repositories, 78-week window
Grouped columns show that commits per engineer rose 14.9 percent year over year and fell 11.2 percent quarter over quarter, while performance per commit rose 100.4 percent year over year and 32.4 percent quarter over quarter, highlighting the change in the value of each commit rather than the number of commits.
Commits per engineer moved 14.9% year over year. Performance per commit moved 100.4%. Quarter over quarter, commits per engineer fell by 11.2%, while performance per commit rose by 32.4%.
Run that comparison on a team where someone is reporting commit volume as a productivity signal, and you’ll see how the story inverts. The same period reads as a slowdown in activity and a large gain in value.
Two things to carry with those figures. The headline quarter-over-quarter figure for performance per engineer came in at +17.5%, with a 95% confidence interval of -0.7% to +38.8%, and the report is clear about what that means: “The quarter-over-quarter change cannot be distinguished from zero.” And the study measures only public repositories, without attributing the shift to assistant adoption. Correlation is what any measurement of this kind gives you, including yours.
Cycle time and change size have to be read together
Third reading, and the one most likely to produce a false positive.
Cycle time falls when changes get smaller. Assistants encourage smaller changes, because a suggestion that fits in one screen is easier to accept than one that spans four files. So cycle time drops, the dashboard turns green, and the quarterly report claims a speedup that is mostly a change in how work was cut up.
Read the two series side by side. Cycle time down 30%, median change size down 40%, and throughput flat means the work got chopped, and the review queue absorbed the difference. Cycle time down 30% with change size steady and throughput up means something real happened.
If you want the cycle time and change size series read against each other on your own repositories, connect a repo, and both come back in the first pass alongside the historical window.
Where the capacity went, by type of work
Fourth reading. This is the one that answers the planning question rather than the finance question, and it’s the one engineering managers tend to find most useful about themselves.
Score each merged change by what it was, then look at the shares. In the Q2 2026 window (Q1 2025 to Q2 2026), maintenance lost 10.9 percentage points of share, while tests gained 8.0 and features gained 4.4. Fixes moved down 1.3 points and docs moved down 0.2 points. Public repositories only, with no visibility into private work, code review depth, incident response, planning, or mentorship.
A team that moved 10 points of its quarter from maintenance into tests and features got faster in a way no velocity chart will show, because velocity counted both categories as points and treated them as equal.
Assistant acceptance rate does not answer this question
Vendors report it because it’s the metric they can see. It tells you a developer pressed tab.
It says nothing about whether the suggestion survived review, whether the resulting change shipped, or whether it introduced the bug that consumed Thursday. A team with 40% acceptance and careful review can be delivering considerably more value than a team with 70% acceptance and a growing revert rate.
Acceptance rate is a usage signal, and usage signals belong in the rollout conversation. Delivery questions need delivery data.
How to report the answer, including when it is no
Write the answer as three sentences and one interval.
What the readings say, over what window, against which baseline. Where the uncertainty is, stated as a range rather than implied by a rounded percentage. And what you’d need to see next quarter to change the conclusion.
Sometimes the honest answer is that throughput rose and you can’t separate the assistant from the two senior hires who joined in the same quarter. Say that. A number with a stated limit reads as instrumentation, and the engineering leader who reports one gets asked harder questions and keeps more credibility.
Then keep the method fixed. The value of this measurement compounds across quarters, and it resets to zero the moment somebody changes the definition to make a quarter look better. Talk to us if you want the four readings run against your own history.
Frequently asked questions
- How long after adopting an AI assistant can I measure the effect?
- Give it a full quarter of normal work after the first few weeks of experimentation, so roughly four months. Earlier than that, you’re measuring the learning curve, which looks like a slowdown followed by a spike, neither of which reflects the steady state.
- Is cycle time a reliable measure of AI speedup?
- On its own, no. Cycle time falls when changes get smaller, which is exactly what assistants encourage, so it can improve while delivered throughput stays flat. Read it alongside median change size and a value-scored throughput figure.
- Should I use developer survey results to answer this?
- Use them alongside the delivery readings rather than in place of them. Perceived friction is real information about retention and daily experience, and it moves independently of what makes it to production. Reporting a sentiment figure as a throughput gain is what gets a productivity claim walked back.
- Does this measurement report on individual developers?
- No. All four readings aggregate at the team and repository levels, where delivery commitments sit. A per-person view becomes a ranking, and that ranking changes the behavior it observes over two sprints.
- What if throughput went up but I cannot prove the assistant caused it?
- That’s the normal result, and stating it is the stronger position. Navigara’s own Q2 2026 study, across 65 public repositories at six organizations, declines to quantify what share of its level shift is attributable to assistant adoption. Report the change, name the other variables in the quarter, and let the decision account for both.
- Can acceptance rate ever be useful?
- As a rollout signal, yes. Near-zero acceptance in one team usually indicates a configuration issue or a language the model handles poorly, both of which are worth fixing. As evidence of delivery speed, it fails because it stops measuring at the keystroke.

