Blog

How to measure engineering throughput without tracking individual developers

Navigara6 min read
Developers working together on software.
Developers working together on software. Photo by Mizuno K · Pexels License

Measure at team and repository level, score each merged change by the value of what shipped rather than by activity counts, and keep per-developer views out of the product entirely. Throughput belongs to a delivery unit. A team owns a commitment. A repository holds a codebase. An individual owns neither of those things, so a per-person number answers a smaller question and costs considerably more to ask.

Why per-developer metrics break the thing they measure

Picture the meeting where it goes wrong. An engineering director opens a tool that ranks fourteen engineers by commits per week. The name at the bottom belongs to the person who spent the quarter untangling the authentication service that everything else depends on, resulting in four commits and saving the company from a November outage. The name at the top belongs to whoever has been renaming variables.

Nobody in that room is behaving badly. The measurement handed them a ranking, and a ranking is very hard to unsee.

Within two sprints, the team has adapted. Pull requests get smaller because small pull requests merge faster. The hard refactor gets deferred because it doesn’t score. Somebody starts splitting a single change across three commits. The numbers improve, and the delivery doesn’t, and the org now has an instrument that reports the opposite of the truth with total confidence.

This is the structural problem with individual measurement, and it has nothing to do with trust. Engineering work distributes unevenly by design. One person unblocks four others. One review prevents three incidents. Pair work produces one commit and two contributions. Every one of those is a team result wearing one person’s name in the metadata.

What a team-level throughput measurement actually reads

The unit of measurement is the merged change, and the question asked of it is what the change was worth, which is a different question from how large it was.

Engineering Throughput Value scores each merged change on what it did: a complex refactor, a performance fix, a feature unblock, a dependency bump, a test backfill. Three lines of code can be any of those. An activity metric shows three lines and then stops.

Aggregate that across a team over a quarter, and you get something a budget meeting can read. Here is the shape of it, from Navigara’s Q2 2026 study, which measured 699 engineers and 137,592 qualifying commits across 65 public repositories at six organizations:

Work mix shift, Q1 2025 to Q2 2026

CategoryChange in share
Tests+8.0pp
Features+4.4pp
Docs-0.2pp
Fixes-1.3pp
Maintenance-10.9pp

Percentage point change in share of work, 78-week window

A stacked bar showing how the mix of engineering work changed across Navigara’s 78-week Q2 2026 window, with tests gaining 8 percentage points of share and features gaining 4.4 percentage points, while maintenance lost 10.9 percentage points and fixes and docs moved slightly down.

Maintenance lost 10.9 points of share while tests gained 8.0 and features gained 4.4. That’s a sentence about where the capacity went, and there’s no individual anywhere in it. It’s also the sentence a CFO asks for, and the one a per-developer leaderboard can’t produce, because a leaderboard measures people against each other rather than measuring work against last year.

What a team-level measurement deliberately ignores

Worth being explicit about, because this is the first question a team asks when they hear a measurement tool is being installed.

Keystrokes. Hours in the editor. Time to first commit of the day. Assistant acceptance rate. Screen activity. Location. Anything that reports on a person rather than on a change that reached production.

The reason is practical before it’s ethical. Every one of those signals is easy to collect and easy to game, and the gaming starts the moment the signal becomes visible. A measurement that changes behavior as it observes it ceases to be a measurement.

How to set this up without starting a monitoring conversation

Four decisions, made before anyone connects anything.

Pick the unit and say it out loud. Team and repository. Write it in the announcement. The announcement is the whole conversation, and an announcement that leads with “visibility into engineering” reads as surveillance no matter what the product does.

Show the team the data first. Before leadership sees a number, the team that produced it sees it. This single sequencing choice determines how the tool is received. Data that arrives downward is an audit. Data the team already has is an instrument.

Name the audience. The number exists to answer questions from finance, the board, and the planning cycle. It does not enter a performance review. If it might, say so now rather than discovering it in April.

Keep the historical window the same. Measuring this quarter against a baseline built on a different set of repositories yields a delta that’s mostly due to scope changes.

If you want to see what the team-level read looks like against your own history before committing to a rollout, connect a repository and the historical window scores in the first pass.

What to do when someone asks for the per-person view

Somebody will. Usually a well-meaning VP who wants to reward the quiet high performer, occasionally a finance partner who has been asked for a productivity number and reasonably assumes productivity is a per-head figure.

The answer that works is specific. Per-developer commit and volume data already exist in Git and have existed for 20 years. Anyone who wanted a leaderboard could have built one on a Saturday. The reason engineering orgs don’t is that the leaderboard is wrong in a way that costs money: it penalizes refactoring, rewards churn, and produces a ranking that inverts under any serious review of what the team actually shipped.

Then offer the thing they were reaching for. A VP who wants to recognize the authentication-service refactor can now do it, because a value-scored change history makes that refactor legible as the largest piece of work in the quarter. The per-person number never showed that. The change-level score does so without ranking anyone.

How team-level data still answers a performance question

Here’s what the team-level read supports, which covers most of what anyone actually needs.

Whether this quarter’s throughput sits above or below the same team’s pre-AI baseline. Where capacity went, split by features, tests, fixes, docs, and maintenance. Which repositories absorbed disproportionate maintenance. Whether cycle time fell because work got faster or smaller. What the AI tooling returned against what it cost.

That’s a defensible case for headcount, a defensible answer on tool renewal, and a defensible board slide. None of it requires knowing what any individual did on Tuesday.

The engineering manager always knew the team was performing. Nothing else in the building could read what they knew. The instrument that finally reads it doesn’t need to point at a person to work.

Talk to us if you want the team-level read run against your own repositories.

Frequently asked questions

Can you accurately measure engineering throughput at the team level?
Yes, and team level is the more accurate unit. Engineering work distributes across people by design, so a per-person split assigns shared results to whoever happens to hold the commit metadata. The team is the unit that owns a delivery commitment, so it’s the unit where the number means something.
Does Navigara report on individual developers?
Measurement runs at team and repository level. The change-level score identifies what a piece of work was, which makes a large refactor visible as a large piece of work without ranking the people who did it.
How do I tell my team we are installing a measurement tool?
Name the unit as team and repository, name the audience as finance and planning, confirm it does not enter performance reviews, and show the team their own data before leadership sees it. The sequencing matters more than the wording.
Is commit count ever a useful metric?
As a data-quality check, yes. A repository with zero commits in a window usually means a broken integration rather than an idle team. As a performance signal, it fails, because it treats a dependency bump and a database migration as the same event.
What about DORA metrics? Are those individual or team level?
DORA is team-level by design and worth keeping. It measures the delivery pipeline rather than the value that moves through it, so it pairs with a throughput measure rather than replacing one.
Can a manager use team-level data to identify a struggling engineer?
Not from the throughput figures, which aggregate above the individual. That conversation runs on code review, pairing, and the manager’s own observations, just as it did before any dashboard existed.

More from the blog