Blog

7 engineering metrics that stop working once AI writes the code

Jirka Bachel6 min read
Vintage computer and keyboard.
Vintage computer and keyboard. Photo by Pixabay · CC0 1.0

Seven metrics degrade once the assistant is written into the code: commit count, lines of code, pull request count, story points, velocity, cycle time, and assistant acceptance rate. Each one breaks for its own reason, and the reasons matter more than the list, because two of these metrics will keep improving on your dashboard while delivery stays exactly where it was.

Commit count

Commit count was already a weak signal. Assistants push it past useful.

The mechanism is decomposition. A developer working with a suggestion loop commits in smaller pieces, because the natural unit of work becomes the accepted suggestion rather than the finished thought. Agentic workflows go further and commit per step. Add generated scaffolding, tests, and configuration files, and the count rises without a single additional delivery.

The Q2 2026 study puts a figure on how little that count carries. Across 699 engineers and 137,592 qualifying commits, commits per engineer increased 14.9% year over year, while performance per commit increased 100.4%.

What a commit was worth, year over year

ComponentYear over year
Commits per engineer+14.9%
Performance per commit+100.4%

Q2 2025 to Q2 2026, 699 engineers, 65 public repositories

Two horizontal bars comparing the components of Navigara’s Q2 2026 year-over-year result, with commits per engineer up 14.9 percent against performance per commit up 100.4 percent, showing that counting commits would have missed almost all of the change.

A dashboard counting commits over that window would have reported a modest year. The value of what each commit carried doubled.

Lines of code

Lines of code were used to measure effort, on the assumption that typing was the expensive part. Generation removes that assumption.

Assistants produce verbose code by default: explicit error handling on every branch, full docstrings, test files that enumerate cases a human would have parameterized. All of it defensible, all of it inflating the count. Meanwhile the highest-value change of the quarter might be a 4-line fix to a connection pool that stopped a recurring incident.

The metric now moves inversely with skill in some cases, because an experienced engineer prunes what the assistant generated, resulting in a smaller diff.

Pull request count

Pull request count inherits the commit-count problem and adds one of its own.

Smaller changes merge faster, so teams learn to open more, smaller pull requests. Then automation starts opening them directly: dependency upgrades, generated client updates, agent-produced fixes for lint failures. Every one is a legitimate pull request. None of them represent a decision anybody made about the product.

Counting them together yields a number that increases with automation coverage. Benchmarking that number against another company compares their automation to yours.

Story points

Points were an estimate of effort, calibrated by a team against its own recent experience. That calibration is what assistants break.

The same ticket that took 5 points in 2022 now involves the same design thinking, the same review, the same deployment risk, and considerably less typing. So teams re-estimate. Some drop the ticket to 3. Some keep it at 5 because the hard part never was the typing. Both are reasonable, and the two teams are now using different currencies.

Points were the measure I defended longest as a former CTO, on the grounds that they were at least honest about being an estimate. The honesty holds. Comparability across the adoption boundary doesn’t exist, and that comparison is what leadership was using them for.

Velocity

Velocity is the sum of points, so it inherits every point problem and compounds them with re-estimation.

There’s a second mechanism specific to velocity. It treats all completed work as equivalent. A quarter of maintenance and a quarter of feature delivery produce the same chart if the point totals match, which matters more now because the mix is moving. Across the Q2 2026 window, from Q1 2025 to Q2 2026, maintenance lost 10.9 percentage points of share while tests gained 8.0, and features gained 4.4. Public repositories only, with no view into private work, code review depth, incident response, planning, or mentorship.

A team that shifted 10 points of its capacity from maintenance to features did something significant. Velocity reported a flat line.

If you want to see what your own work mix and value per change did across the adoption boundary, connect a repository and the historical window scores in the first pass.

Cycle time

Cycle time is the strongest case on this list because it stays green even as the thing it represents gets worse.

The mechanism is direct. Cycle time falls when changes get smaller. Assistants encourage smaller changes, because a suggestion that fits on one screen is easier to accept, verify, and ship than one that spans files. So the median change shrinks, the clock per change shrinks with it, and the dashboard shows a team that got 30% faster.

What happened is that the same work arrived in three pieces instead of one. Total elapsed time from ticket to production may be unchanged or longer, because each piece needs its own review, its own CI run and its deployment review queue absorbs the difference, which shows up as senior engineer time and never as a metric.

Cycle time is worth keeping as a pipeline measure. Read it next to the median change size and a value-scored throughput figure, or the trend will tell you the opposite of the truth with a clean line and a green arrow.

Assistant acceptance rate

This one arrived with the tools, which is why you should be careful with it.

Acceptance rate measures how often a developer pressed Tab. It doesn’t distinguish a suggestion that shipped unchanged from one that was accepted, rewritten twice, and reverted on Thursday. It counts the accept and stops.

It also moves for reasons unconnected to delivery. Language coverage, codebase age, how idiomatic the internal frameworks are, whether the team writes in a house style the model has seen. A platform team working in a 12-year-old internal framework will show low acceptance and may be the most productive group in the building.

Use it to find configuration problems during a rollout. A team at near-zero acceptance usually has something broken. Past that point it’s a usage figure, and usage figures can’t answer a delivery question.

What still reads correctly once assistants are in the loop

One measurement survives all seven mechanisms above: the value of what merged, scored per change, aggregated to the repository, and compared against the same team’s own history.

Navigara scores each merged change by what it was rather than how big it was, which is why the decomposition above is possible at all. A complex refactor, a performance fix, a feature unblock, a dependency bump, and a test backfill are five different events. Three lines of code can be any of them, and every metric on this list treats them as identical.

That measurement stays deliberately at team level; per-person activity data has existed in Git for 20 years, and anyone who wanted a leaderboard could have built one already. The reason engineering orgs avoid it is that it penalizes refactoring, which is the same failure mode as counting commits, aimed at a person.

The engineering manager in the room already knows which work mattered this quarter. Talk to us if you want the instrument that reports it in a form the budget meeting can read.

Frequently asked questions

Should we stop tracking cycle time entirely?
Keep it, and read it with change. Cycle time is a real measure of how quickly a change moves through the pipeline, but it becomes misleading when the size of the change shifts underneath it. Falling cycle time plus falling change size plus flat throughput means the work got divided, not sped up.
Are DORA metrics on this list?
No. DORA measures the delivery pipeline: deployment frequency, lead time for changes, change failure rate, and time to restore service. Those still work and are worth keeping. They describe how change moves rather than what the change was worth, so they pair with a throughput measure.
What replaces story points for planning?
Nothing needs to, for planning. Points work fine as a team’s internal estimate for the next two weeks. The failure is using them for cross-quarter or cross-team performance comparisons, because the calibration behind them changed when the tools did.
Why do lines of code get worse rather than just staying weak?
Because generation makes the count cheap to increase, and the count was always assumed to track effort. An engineer who prunes generated code lands a smaller diff for more work, so the metric can now move in the opposite direction to the contribution.
Does measuring value per change require monitoring developers?
No. The unit of measurement is the merged change, and the aggregation is by team and repository. There’s no per-person view, no editor telemetry, and no keystroke data used to score what a change was.
How many quarters of data do I need before this comparison means anything?
Two pre-assistant quarters as a baseline and at least one full quarter after adoption settled. A single quarter on either side carries too much noise from one large release, and the Q2 2026 study is explicit that “A single quarter of deceleration in a noisy series is not yet evidence of a plateau.”

More from the blog