Blog

Nobody Will Work Without AI Now. Nobody Can Estimate Anything Either.

Jirka Bachel and Peter Malina12 min read

Twelve months ago, engineers came to meetups to tell us the research proved AI was slowing them down. They were right about the research. In 2025 METR ran a randomized trial on experienced open-source developers and measured them 19% slower with AI, while those same developers estimated afterwards that they had been 20% faster [1].

METR revisited it this year. For a subset of the original cohort they now estimate roughly an 18% speedup [2].

Worth being precise about how much that second number carries, because it is weaker than the headline suggests. The confidence interval runs from -38% to +9%, so it crosses zero, and it covers only part of the original group. The direction moved. The certainty did not.

The finding underneath it is the one that actually matters. METR had to change the experiment design, because they could not recruit enough developers willing to work without AI. Between 30 and 50% of participants were holding tasks back rather than do them unassisted. In a year we went from engineers arguing that the data proved AI slows you down, to being unable to assemble a control group.

Everybody says it is working, and nobody can say so honestly

Every company is running the survey now. Ask developers whether AI is helping and the answer comes back yes, overwhelmingly, everywhere.

That result is close to worthless, and not because developers are lying. It has become genuinely difficult to say out loud that AI is not helping you. You need to be a confident person with your arguments ready, in a room where the budget has already been approved and the strategy already announced. Saying you use it for everything is free. Saying it slows you down costs something.

The more textured surveys show the shape of it. Roughly 70% report AI is helping, and 40 to 45% report it is helping but is not there yet, which is the frustrating middle. You get 90% of the way and the last stretch is maddening, because the tool did a pile of things nobody asked for.

So the honest position is that we have moved the bottleneck rather than removed it. Writing code is no longer the constraint. Reviewing it is, and deciding what to build is. Almost all output measurement sits on pull requests, and so does almost all review time.

The layoffs are not about speed

Here is the part that gets skipped in every discussion about AI and headcount.

The reason companies put AI into engineering was to lower the cost of engineering. Keep the same headcount and add AI spend on top, and engineering costs went up. Not down.

Now follow the logic one step further. If you knew that delivering your roadmap faster would increase revenue, you would already be hiring engineers. That is not a hard decision, and it never was. The reason a company is not hiring is that it is not confident faster delivery produces revenue.

Which means most of the layoffs being attributed to AI are not because AI made teams faster. They are because the budget blew up and nobody could show what it bought. Those are very different problems with very different fixes, and only one of them is solved by removing people.

Do we still need engineers

The question arrives constantly now, because AI is the first one of these tools that is not for engineers. A product manager was never going to open an IDE. Now they build something that looks like it works over a weekend and ask why twenty people need three months.

We think the premise is wrong, and the reason is unglamorous. Every department in the company is onboarding these tools right now. Every department is being automated as we speak. Somebody has to onboard those tools, improve them, wire them together, and answer the pager when one fails. That is engineering work. HR is not going to build a workflow that reaches into every system in the company.

The role changes, which is the part people mistake for the role disappearing. There will be more engineers, because the future is more technical than the present, not less. Jevons paradox holds here as it does everywhere: the easier something gets, the more of it gets used. And if engineering work could genuinely be fully automated, you would not need CEOs either. You would not need companies.

What the weekend prototype leaves out is most of the work. Going from prototype to production has always been the expensive part, sometimes 50 or 100 times the effort once security, performance, infrastructure, and scale are real. The edge cases are not a detail around the happy path. They are 80% of the product.

Apiiro put numbers on the cost of skipping that. Studying a single Fortune 20 enterprise, more than 7,000 developers across 62,000 repositories, they found AI-assisted developers producing three to four times more commits and ten times more security issues. Those developers exposed cloud credentials and keys nearly twice as often, and their work arrived in fewer, larger pull requests, which is precisely the shape that produces shallow review [3].

Two bars against a 1x baseline: AI-assisted developers produced 3 to 4 times more commits, and 10 times more security issues.
Both bars are measured against the same developer without AI. Output roughly tripled. Defects went up an order of magnitude.

Three to four times the output with ten times the vulnerabilities is not a result you can take to a board. The checks that catch this can be built, and built with AI, but running them fully automated is expensive enough that hiring a person is sometimes the cheaper answer.

The ceiling arrived earlier than expected

Our open source performance index showed enormous jumps through Q1. In Q2 it showed nothing. Performance per engineer was flat across the entire quarter.

Flat, not falling. The teams at the bottom are still ramping, because they started slow. The teams at the top appear to have hit something. The spread we measure is still 30 to 50 times what it was before AI, so the ceiling is high. It is still a ceiling.

Line chart of performance per engineer indexed to Q1 2025, rising +0%, +9%, +18%, +25%, then jumping to +116% in Q1 2026 and staying flat through Q2 2026.
Four quarters of steady single-digit gains, one 91-point jump, then nothing. The Q2 flat line is the part we did not expect this early.

The best explanation we have is absorption. You cannot take more complexity into your head than you can hold, and you should not ship what you are not willing to answer for. Generation is no longer the limit. Judgement is.

There is a related pattern in the data that surprised us. Introducing a coding agent to a team looks, in the numbers, like adding new people to that team. Performance drops first, for a month or two, because the way the team builds is changing and everyone has to relearn it. Nobody expects a new licence to lower output. It is not in the marketing.

Team size behaves the same way it always did, only more visibly. Average performance per engineer falls as the team grows. One person has very high per-person output. Add five, twenty, fifty, and the average drops. That is the real argument hiding inside all the talk about engineers needing to understand product and business: it is an argument for smaller teams, because shrinking the team raises the average.

The talent gap was already there

The hope was that AI would compress the distance between a junior and a senior. It did the opposite, and GitClear has the cleanest measurement of why.

Heavy AI users out-produce non-users by four to ten times. But most of that gap pre-dated the tools. Measured against their own output twelve months earlier, those same heavy users gained about 25% [4].

Two horizontal bars: heavy AI users out-produce non-users by 4 to 10 times, but only out-produce their own output from a year earlier by 1.25 times.
The same people, measured two ways. Against their peers the gap is enormous. Against themselves a year ago it is 25%.

So the tools did not create high performers and they did not close the gap. High performers adopted them first. Heavy AI usage is a signal that somebody was already strong, not an explanation of why.

Our own data says the same thing about distribution. Roughly 8% of the engineers we measure deliver more than half of the output. The chart is boring until you read the axis.

The scale underneath that: average ETV per engineer per month sits around 2. Fifty a month is the top of the top, and we struggle to find anyone above a hundred. The engineer producing fifty is not paid twenty-five times the average. We expect the spread in engineering salaries to widen a great deal.

What separates those people is not technical depth. Every engineer is motivated to learn more technology. Far fewer learn the business. An engineer who has spoken to the customer can make most decisions alone while implementing. An engineer who has not builds something, brings it to review, gets told it was the wrong thing, and goes back around. Best case that feedback comes from a product manager in the same week rather than a director three weeks later. That loop is where the months go.

Every instrument we have left is self-reported

We are not fans of most engineering metrics, because most of them are wrong in a way that redirects effort rather than measuring it.

Velocity is the common one, because it sits next to Jira and feels like roadmap progress. It runs entirely on self-estimates. How do you prove you are exceptional? Estimate high, deliver fast. Anyone who has run an agency has watched this happen: you delivered 40 points, now do 45, and you can see the room recalibrating. What was a three becomes a five. Two weeks later the team delivers 50 points and the customer points out that these are the same features as last time. The customer is right.

Estimating agent work makes it worse. Any given task takes two hours or two weeks and there is little warning which. Whatever number goes into the meeting, you find out afterwards.

Cycle time fails from the other direction. Tickets are cheap to create with AI, so smaller tickets produce shorter cycle times and no extra delivery. Make AI spend the metric and the engineer who avoids AI becomes your best performer. Count pull requests and a staff engineer who shipped one change through security review and QA looks idle next to a junior who went back and forth forty times.

This is why we built ETV. Delivering the roadmap is why engineers were hired, and judging that requires knowing how hard the work actually was. You cannot get that from a self-estimate, and the people who deserve a bigger budget need something they can put in front of a CFO. Your CEO does not care about cycle time; you have to explain cycle time to him first. Your CFO has no idea what it means, and he is the one approving the spend.

Speed is cheap now. Alignment is the whole game.

The most uncomfortable pattern in our customer data: the highest performers in almost every instance are the people contributing least to the roadmap.

It is easy to look fast when you answer only to yourself. You build what your own tooling needs and ignore where the company is going. Everyone working on roadmap items is two to three times slower, because they have to listen to customers. Then ask which of those two had more impact, and the business answers immediately.

The same dynamic shows up in spend. If someone is delivering roadmap work that increases revenue, that is the person who should be spending five or ten thousand dollars a month on tokens, and you should be pleased about it. Instead there is usually a ceiling of a few hundred dollars, and the message an engineer takes from a ceiling is that a good employee does not complain. So they go and find a cheaper model with a better harness. They optimise the cap instead of the roadmap, which is the correct response to being handed a target. We have written about why caps backfire.

The other failure mode is nobody watching the denominator at all. One company was spending about $2,000 to run a single CI/CD pipeline, with a junior developer pushing twenty pull requests a day into it. Four thousand dollars a day to review one junior’s work. Nobody set out to do that. Somebody wanted the pipeline safe enough that a non-technical person could contribute to enterprise software, which is a reasonable goal that arrived with an unreasonable bill.

So the gap we most want to close is alignment. One engineer can reach a certain level of performance. Why can twenty not reach the same number between them? That is not a tooling problem.

Nobody can estimate anymore, and capacity is what is left

We struggle to find a founder using coding agents who will claim, with confidence, that they can still estimate. Estimation ran on having done something similar before, in that part of the codebase. The model changes every month, everyone is still learning to prompt, and there is no benchmark to sit against. You can say five story points. Based on what?

Every sufficiently complex project used to reach the point where honest estimation became impossible. Now most work reaches it immediately.

What you can still know is where the capacity went. Not how long an initiative will take, but how much of the team’s delivered output landed on the initiatives that matter rather than on keeping the lights on. Whether something takes one month or two is often immaterial, because both are acceptable to your customers. Whether anyone was working on the right thing is never immaterial.

That is the measurement we think survives this. Speed stopped being the scarce thing. Knowing what to point it at did not.

References

[1] METR, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity,” 2025. [Online]. Available: https://metr.org/blog/2025-07-10-early-2025-ai-experienced-os-dev-study/

[2] METR, “We are Changing our Developer Productivity Experiment Design,” February 2026. [Online]. Available: https://metr.org/blog/2026-02-24-uplift-update/

[3] Apiiro, “Faster code, greater risks: The security trade-off of AI-driven development,” September 2025. [Online]. Available: https://apiiro.com/blog/faster-code-greater-risks-the-security-trade-off-of-ai-driven-development/

[4] GitClear, “AI Coding Tools Attract Top Performers, But Do They Create Them?” 2026. [Online]. Available: https://www.gitclear.com/developer_ai_productivity_analysis_tools_research_2026

[5] Navigara, “Methodology: How Navigara Measures Engineering Throughput,” 2026. [Online]. Available: https://research.navigara.com/methodology

Frequently asked questions

Did the METR study reverse itself on AI making developers slower?
Partly. METR's 2025 randomized trial measured experienced open-source developers as 19% slower with AI, while those same developers estimated they had been 20% faster. A 2026 follow-up on a subset of that cohort estimates roughly an 18% speedup, but with a confidence interval from -38% to +9%, so it crosses zero. The direction moved; the certainty did not. The more decisive finding is that METR had to redesign the experiment because developers would no longer take on tasks without AI.
Does AI close the gap between junior and senior engineers?
The evidence points the other way. GitClear found heavy AI users out-produce non-users by four to ten times, but most of that gap pre-dated AI: measured against their own output twelve months earlier, those same heavy users gained about 25%. The tools were adopted first by people who were already high performers, so they amplify an existing gap rather than closing it.
Why don't velocity and story points work for measuring AI-assisted work?
Both run on self-estimates, and estimates recalibrate the moment they become a target. Ask a team for 45 points instead of 40 and the same work gets priced at 45. Cycle time has the same flaw from the other direction: tickets are cheap to create with AI, so smaller tickets produce shorter cycle times and no additional delivery. Scoring the actual complexity of shipped code, as ETV does, is the only version that stays comparable over time.
Can you still estimate engineering work when the team uses coding agents?
Mostly no, and it is worth saying so plainly. Estimation relied on having done something similar before in the same codebase, and with agents any given task takes two hours or two weeks with little warning which. The next best measure is capacity: how much of the team's effort is landing on initiatives that matter versus keeping the lights on. Whether an initiative takes one month or two is often immaterial if both are acceptable to customers.

More from the blog