Why DORA metrics broke after AI, and what replaced them

DORA measures the delivery pipeline, and it measures it the same way for AI as for before: deployment frequency, lead time for changes, change failure rate, and time to restore service. What changed is the weight teams now put on it. Four green DORA metrics say change moves through your pipeline quickly and safely. They say nothing about the worth of what moved, and that’s the question the AI budget conversation turns on.
What the four metrics measure, precisely
Deployment frequency: how often you release to production. Lead time for changes: how long it takes for a commit to reach production. Change failure rate: the share of deployments that cause degradation requiring remediation. Time to restore service: how long recovery takes when one does.
Read that list again and notice what all four have in common. Every one of them describes the movement of change through a system. Speed and safety of transit.
That was the point. DORA emerged from research into delivery capability, and as a measure of it, it holds up better than almost anything else engineering has produced. Teams that improve those four numbers do ship more reliably. The research behind it stands.
Why the four metrics still hold after AI
Worth saying plainly, because plenty of commentary this year suggests otherwise.
An assistant writing part of the code does not change any of the definitions. A deployment is still a deployment. A change that breaks production still counts as a change failure whether a human or a model wrote the line. Time to restore service is measured in minutes of customer impact, and models don’t alter the arithmetic of an outage.
If anything, the change failure rate and time to restore became more load-bearing because the volume of change entering the pipeline increased. A team merging 40% more changes per week while maintaining a stable failure rate has real evidence of its review and testing discipline. Keep all four. Instrument them properly. They’re the cheapest early warning you have.
What the four metrics leave out
Here’s the structural gap, and it predates AI by a decade.
DORA counts deployments. It doesn’t ask what was in them. A quarter of 900 dependency bumps and a quarter containing the authentication refactor that prevented a November outage can produce identical deployment frequency, identical lead time, identical failure rate. The pipeline performed the same. The company did not.
Picture the board prep meeting. A platform lead has four green metrics on a slide, deployment frequency up 60% since last year, and a director asking what the company got for a year of assistant licenses and token spend. The four metrics have no answer available. That variable sits outside everything they measure.
Before AI, that gap was tolerable. Change volume was roughly proportional to work, so counting deployments was a rough proxy for delivery. Assistants removed the proportionality. Volume now rises with tooling coverage, with how work gets cut up, and with automation opening pull requests nobody planned.
What AI did to the reading of each one
Each of the four now needs a second series next to it.
Deployment frequency rises when changes get smaller, which assistants encourage. Read it with median change size, or the trend reports a throughput gain that is just a chopping up of the same work.
Lead time for changes falls for the same reason, and it can fall while total ticket-to-production time rises, because one change became three, each with its own review and CI run.
Change failure rate is the one that can hide a real problem. Generated code tends to fail in ways that pass tests and surface later: a subtle concurrency assumption, an unhandled edge in an error path. Pair it with revert rate and with time-to-detect.
Time to restore service stays clean, which makes it the most reliable of the four right now.
The companion measure that closes the gap
The missing quantity is the value of each merged change, scored by what the change was rather than how large it was or how fast it traveled.
Navigara calls this Engineering Throughput Value. Each merged change is scored on what it did: a complex refactor, a performance fix, a feature unblock, a dependency bump, a test backfill. Five different events, and three lines of code can be any of them.
The Q2 2026 study shows why the pairing matters. Across 699 engineers and 137,592 qualifying commits in 65 public repositories, commits per engineer increased by 14.9% year over year, while performance per commit increased by 100.4%. Activity, which is what a pipeline metric can see, barely moved. The worth of each change doubled. Public repositories only, and the study doesn’t quantify what share of that level shift, if any, is attributable to the adoption of AI coding assistants.
The same study also reads the composition of the work, which no DORA metric can.
Change in share of work, Q1 2025 to Q2 2026
| Category | Change in share |
|---|---|
| Tests | +8.0pp |
| Features | +4.4pp |
| Docs | -0.2pp |
| Fixes | -1.3pp |
| Maintenance | -10.9pp |
Percentage point change across a 78-week window
A stacked bar showing how the composition of engineering work shifted over the 78-week window, with tests gaining 8 percentage points of share and features gaining 4.4 percentage points, while maintenance lost 10.9 percentage points and fixes and docs moved slightly down.
Maintenance lost 10.9 points of share while tests gained 8.0 and features gained 4.4. Deployment frequency cannot express that sentence, and it’s the sentence a planning cycle needs.
If you want the value-per-change read sitting next to the DORA figures you already collect, connect a repository and the first pass scores your historical window too.
How to report both together without doubling the slide
One slide, two halves, and a rule about which half answers which question.
The pipeline half carries the four DORA metrics with their second series: deployment frequency with change size; lead time with ticket-to-production time; change failure rate with revert rate; and time to restore on its own. That half answers whether delivery is healthy.
The value half carries throughput scored per change, measured against the team’s own pre-AI quarters, plus the work mix split. That half answers what the quarter was worth and what the tooling returned.
State the uncertainty on both. The Q2 2026 study is a useful model here: its quarter-over-quarter figure came with a 95% confidence interval from -0.7% to +38.8% and the flat statement that “The quarter-over-quarter change cannot be distinguished from zero.” A number carrying its own interval reads as instrumentation. I reported plenty of numbers as a former CTO that would have survived longer if I’d published the interval with them.
Keep the value half at team and repository level. Per-person throughput reporting turns the measurement into a ranking, and a ranking changes the behavior it observes inside two sprints. The team is the unit that owns a delivery commitment, so it’s the unit where the figure means something.
What to say when somebody proposes dropping DORA
Someone will, usually after reading that the metrics no longer apply in an AI era. The argument to make is specific.
The four metrics are the only instruments on the slide that will tell you when the pipeline is failing, and the volume of change hitting that pipeline has gone up. Dropping them at this exact moment removes the smoke detector during the renovation.
What the four need is company. Delivery health from DORA and delivered value from a throughput measure, scored per change, both read against your own history rather than an industry average. That combination answers the finance question, the planning question, and the incident question without any of the three borrowing evidence from the others.
Talk to us if you want the value half built against your repositories while you keep the pipeline half where it is.
Frequently asked questions
- Are DORA metrics still valid in 2026?
- Yes. The four metrics measure the delivery pipeline, and their definitions are unaffected by who or what wrote the code. They remain the best available read on delivery health, and they say nothing about the value of what shipped, which is why most teams now pair them with a throughput measure.
- Which DORA metric is most affected by AI coding assistants?
- Deployment frequency and lead time for changes, because both move when the size of a change moves. Assistants encourage smaller changes, so both metrics improve while the same amount of work is being delivered in more pieces. Read each with a change-size series.
- Can DORA metrics show the ROI of AI coding tools?
- Not on their own. They measure how fast and how safely change moves through the pipeline, and an ROI question asks about the worth of what moved. Pairing DORA with a value-scored throughput figure against your own pre-AI baseline covers both halves.
- Is change failure rate reliable when AI writes part of the code?
- It’s reliable as a definition and incomplete as a signal. Generated code can fail in ways that pass tests and appear weeks later, so a stable failure rate may sit alongside a rising revert rate or a longer time to detect. Track those two next to it.
- Do we need a different set of metrics for AI-assisted teams?
- You need one additional, not a new set. Keep the four pipeline metrics, add a measure of the value of each merged change, aggregated at the team and repository levels, and compare against your own history. That combination answers delivery health and delivered worth from the same dataset.

