Blog

How do I set a pre-AI baseline for my engineering team?

Jirka Bachel6 min read
Developers reviewing code on a large display.
Developers reviewing code on a large display. Photo by Mikhail Nilov · Pexels License

You set a pre-AI baseline by scoring two full quarters of your own Git history from a period that ended before the first assistant entered the workflow, then keeping that exact method for every measurement you run afterward. The data already exists. Nobody has to instrument anything new. The work sits in choosing the window, cleaning the contributor list, and refusing to measure the past one way and the present another.

What a baseline is actually holding

A baseline is a single figure with three components attached: a window, a repository list, and a method. What this team shipped per quarter, valued by the size and risk of each change, during a period when no assistant was in the loop.

Every claim you make about AI later is a comparison against that figure. Without it, you’re comparing this quarter to last quarter, and both quarters include assistants, so the comparison can’t answer the question anyone is asking.

The question arrives on a schedule. Renewal season, or the week token spend triples because somebody moved the team from a chat-style assistant to an agentic workflow that reads half the codebase before it writes a line. Then a finance partner asks what the first year bought.

I’ve been on the wrong side of that meeting. As a former CTO, I signed off on tooling spend based on how the team felt about it, which is a real signal but a completely unusable one the moment somebody wants the arithmetic.

Why your own history beats an industry benchmark

An industry average doesn’t account for any of the factors that determine your throughput. Your merge policy, your review culture, your test burden, the age of the codebase, how much of the work is public versus private.

Navigara’s own Q2 2026 study makes the spread visible. Across 65 public repositories at six organizations, year-over-year change in performance per engineer ran from +38.0% at Meta to +151.1% at Vercel, with OpenAI at +230% on a baseline rebased to Q3 2025. Same window, same method, same 78 weeks for five of the six. A benchmark that averages that spread tells you almost nothing about any of them. Public repositories only, and the study doesn’t quantify what share of the level shift, if any, is attributable to the adoption of AI coding assistants.

Your history is the only comparison that still holds your own variables.

How to pick the window

Two full quarters. One quarter carries too much noise from a single large release, and three starts pulling in a team that was a different size.

Four rules for the window, and every one of them exists because a team got burned by skipping it:

End it before the first assistant, not before the rollout. Official rollout dates lag real usage by months. For most US teams, the honest cut is somewhere between late 2022 and mid 2023, depending on when Copilot stopped being a curiosity and became the default in the editor. Ask three engineers when they started using it. Take the earliest answer.

Exclude a reorg. If the team changed size by more than about 20% inside the window, the baseline is measuring the reorg.

Exclude a migration. A monolith-to-services quarter produces a volume of change unrelated to feature throughput. So does a framework upgrade that touched 400 files.

Freeze the repository set. A baseline drawn from 4 repositories, compared later against 9, produces a delta that is mostly scope change wearing a percentage sign.

How to clean the contributor list before you score anything

This is the step that eats the two weeks, and skipping it is why most internal attempts produce a number nobody believes.

Bot accounts first. Dependabot, Renovate, release bots, the CI account that commits generated clients. A historical window full of automated dependency bumps will show a baseline far higher than the team ever produced, and the AI-era comparison will look flat as a result.

Then the identity mess. Engineers who changed laptops and got a second Git identity. Contractors who committed under a personal address. Squash-merge policy changes, which reshape commit counts overnight without changing a thing about delivery. Vendored code drops arrive as a single enormous commit and distort any size-weighted measure.

Then the gaps. A three-week hole in the history usually means the CI provider changed and the commit metadata changed with it. Find it before it becomes a finding.

What to score once the window is clean

Count-based measures are the trap here because they’re the easiest to compute from history. Commits, lines changed, pull requests merged. All available, all cheap, all wrong for this purpose.

Score the value of each merged change instead: a complex refactor, a performance fix, a feature unblock, a dependency bump, a test backfill. Three lines of code can be any of those. Navigara does this with Engineering Throughput Value, and the reason it matters for a baseline is in the study’s own decomposition.

Activity versus value per change, year over year

ComponentYear over year
Commits per engineer+14.9%
Performance per commit+100.4%

Q2 2025 to Q2 2026, 699 engineers, 65 public repositories

Two columns compare components of Navigara’s Q2 2026 year-over-year results, showing commits per engineer up 14.9 percent while performance per commit rose 100.4 percent, so almost all of the movement lies in the value of each change rather than in how many changes there were.

Commits per engineer moved 14.9% year over year. Performance per commit moved 100.4%. A baseline built on commit counts would have reported that almost nothing changed, and it would have been wrong by a factor of seven.

If you’d rather see the historical window scored before you build the internal pipeline, connect a repository and the pre-AI quarters come back in the first pass.

What a finished baseline looks like

One page. Team, repository list, the two quarters, the per-quarter throughput figure, and the work mix split across features, tests, fixes, docs, and maintenance. Plus the exclusions, written down, because the first question you’ll get is what you left out.

And a sentence about the method, phrased so that the person who reruns it in 18 months produces a comparable figure. That sentence is the baseline. The number is just this quarter’s reading.

Say the audience out loud while you’re at it. The baseline exists to answer finance, the board, and the planning cycle. It aggregates at team and repository level, so there’s no per-person view to worry about, and the announcement should say that before anyone asks. A measurement that arrives without that sentence gets read as an audit, and a team that thinks it’s being audited changes what it commits.

What to do if the window has already closed

It hasn’t. The history is still in Git, and it doesn’t expire.

Adoption also arrived unevenly, which helps. One team standardized on an assistant in early 2023 while the platform group held out until 2025. That gives you a within-company comparison, which is better than a benchmark or a single company-wide cut.

The first pass will be wrong. A repository nobody remembered, a bot you missed, a quarter where half the team was on a customer escalation. Fix those, rerun, and the second reading is the one that survives the meeting. Talk to us if you’d rather have the first pass run against your history than build the tooling for it.

Frequently asked questions

How far back does a pre-AI baseline need to go?
Two full quarters ending before assistants entered the workflow, which for most US teams means somewhere between late 2022 and mid 2023. One quarter carries too much single-release noise, and going back further starts to measure a team that was a different size, with different repositories.
Can I set a baseline if we adopted AI tools before we started measuring?
Yes. Git history is the record, and it’s complete whether or not anyone was measuring at the time. The constraint is method consistency: the historical window has to be scored the same way you’ll score every quarter afterward.
Does setting a baseline require tracking individual developers?
No. A baseline aggregates at team and repository level, which is where delivery commitments sit anyway. Per-person history has existed in Git for 20 years, and building a leaderboard from it changes behavior within about two sprints without answering the baseline question.
What should I exclude from the baseline window?
Bot and automation accounts, a reorg that moved team size by more than 20%, a large migration or framework upgrade, vendored code drops, and any period when the squash-merge policy changed. Each of those distorts the figure in a direction you can’t unpick later.
Why not use the commit count as the baseline, since it is easy to compute?
Because commit counts barely move while the value of each change does. In the Q2 2026 study, commits per engineer rose 14.9% year over year while performance per commit rose 100.4%, across 65 public repositories at six organizations. Public repositories only, and the study doesn’t attribute that shift to assistant adoption.

More from the blog