10 ways to cut AI token spend without slowing the team down

Most token spend goes on context nobody reads and turns nobody needs, so the reductions that work are mechanical rather than behavioral. Scope what you send, cache what repeats, cap what loops, and route cheap tasks to cheap models. Ten changes follow. Each is something a team can ship in a sprint, and none limits what an engineer is allowed to attempt.
What to measure before you change anything
A reduction you cannot attribute will be reversed in three months when the bill moves again.
Pull the usage export from every provider account, by day and by API key, for the last 60 days. Take the current published per-million-token rate for each model in use and multiply. Then note which keys belong to which workflow, because the answer is almost always concentrated: one repository, one agent configuration, or one week.
We publish no benchmark for reasonable token spend, and any dollar figure we printed here would be invented. Your invoice is the baseline.
Cut the context you send on every turn
Context is the largest line for almost every team that has adopted agents, because it gets re-sent on every turn of every task.
1. Scope context to paths rather than repositories. An agent handed the whole codebase pays for it on every turn. Configure scope at the package or directory level, and make whole-repository context an explicit choice for a specific task. Patch quality usually improves as well because the relevant files no longer compete with 900 irrelevant ones.
2. Turn on prompt caching for the stable prefix. System instructions, coding standards, architecture notes, and schema definitions repeat identically across thousands of requests. Every major provider bills cached prefix tokens at a steep discount. This requires putting the stable content first and the variable content last, which is a template change rather than a workflow change.
3. Cap retrieval results instead of pasting files. Where a workflow pulls code by search, set a hard limit on the number of chunks returned and their size. Twelve well-scored chunks and 200 chunks produce similar answers at wildly different costs, and the 200-chunk version is slower for the engineer waiting on it.
Bound the loops
An unbounded agent loop has no upper cost, which means your forecast has no upper bound either.
4. Set a step cap on every agent workflow. A hard maximum on turns per task, with a clear failure message when it is hit. This converts an open-ended cost into a bounded one and surfaces the tasks that need more room as visible exceptions rather than silent spend.
5. Fix the flaky tests that trigger retry storms. An agent wrapped around an unreliable suite will retry with full context every time, and a single flaky integration test can dominate a month’s consumption. The usage export shows this as a tight cluster of near-identical requests to a single repository. Quarantining that test is often the largest single reduction available.
6. Trigger expensive workflows on merge rather than on every push. Agent-driven review running on all 14 pushes to a branch bills 14 times for a diff that changed twice. Run the cheap checks on push and the expensive pass once on the pull request, with re-runs on request rather than on commit.
Match the model to the task
Teams commonly route everything to the largest available model, including work that a smaller model handles at full quality.
7. Route by task class. Commit messages, renaming, formatting, test scaffolding, docstring generation, and changelog drafting run acceptably on a smaller tier. Design discussions, debugging, migrations, and anything that touches a public interface belong on the larger one. Write the routing table down and put it in the repository, because a rule that lives in one engineer’s head has no effect on the invoice.
8. Submit asynchronous work through batch endpoints. Providers commonly discount batch processing substantially against interactive rates. Test backfills, documentation passes, dependency triage, and bulk code annotation have nobody waiting on them. Interactive pricing for work with no human in the loop is a cost without a corresponding benefit.
9. Summarise state between turns instead of re-sending the transcript. Long agent sessions re-send the entire conversation on every turn, so a 30-turn task pays for turn 1 thirty times. Compacting the history into a short state summary at intervals sharply reduces consumption on long tasks and tends to improve reliability, since the model stops re-reading its own abandoned attempts.
Remove the spend nobody is using
10. Reconcile accounts, keys, and seats quarterly. Provisioned seats nobody logs into, API keys belonging to departed engineers, a second provider account opened during a spike, and a proof-of-concept service still calling a model on a cron schedule. Finance is paying for all four. This is the least interesting item on the list and frequently the fastest money.
If you want the delivery side measured over the same window as the reductions, connect a repository and the historical window scores in the first pass.
How to tell whether a cut slowed the team down
Every item above affects cost and could, in principle, affect delivery, so the measurement has to run on both sides of the change.
The measurement that matters is the value of what shipped, not the count of what shipped. Navigara’s Q2 2026 study, covering 699 engineers and 137,592 qualifying commits across 65 public repositories in 6 organizations, separates the annual change into two components, and that separation is the part worth copying into your own before-and-after.
Components of the Q2 2026 change in performance per engineer
| Component | Year over year | Quarter over quarter |
|---|---|---|
| Commits per engineer | +14.9% | -11.2% |
| Performance per commit | +100.4% | +32.4% |
Year over year and quarter over quarter, 699 engineers, 65 public repositories
Grouped columns show commits per engineer up 14.9 percent year over year, while performance per commit is up 100.4 percent, with quarterly figures of minus 11.2 and plus 32.4 percent, so the annual change is driven more by the value of each change than by the number of changes.
Commits per engineer moved 14.9% year over year. Performance per commit moved 100.4%. A team watching commit volume after a cost reduction is watching the component that barely moves. A team monitoring the value of merged changes will see a slowdown in the quarter it happens.
Two caveats accompany that figure and should be included in any internal deck citing it. The study does not quantify what share of the level shift, if any, is attributable to assistant adoption. And 21 of the 65 repositories in the sample are AI or agent SDKs, unevenly distributed across the 6 organizations, so for those repositories, throughput per engineer and category demand cannot be separated.
The order to do these in
Start with the usage export, because it tells you which of the ten applies to you. Then prompt caching and context scope, which are configuration changes with no behavioral component. Then step caps and the flaky test. Then routing and batch, which need a written rule and a conversation. Account reconciliation last, because it is the one that needs finance in the room.
Nine of the ten are things an engineer can make a pull request for on a Thursday. The tenth needs a meeting, which is why it is still costing you money.
Talk to us if you want cost and delivery read against the same window before you start cutting.
Frequently asked questions
- Which change cuts AI token spend the most?
- Usually context scope, followed by prompt caching, because context gets re-sent on every turn of every task while a seat gets billed once. Your own usage export will tell you which of the two dominates, since the ratio depends on how many turns your typical task takes.
- Does prompt caching change the quality of results?
- No. Caching stores the identical prefix you would have sent anyway and bills it at a lower rate. The model sees the same content, so the only change is the invoice and a modest improvement in latency.
- Are agent step caps going to block real work?
- A cap set near your observed distribution rarely fires, and when it does, it surfaces a task that needs a different approach. Set the limit from your own turn-count data rather than a round number, and treat every cap hit as a signal worth reading.
- How much can we expect to save?
- We publish no figure, and any percentage quoted here would be invented. The honest method is to measure 60 days of current consumption from your provider’s published rates, apply two or three of these changes, and measure the same window again against the same task mix.
- Will cutting token spend slow the team down?
- Scoping, caching, routing, and batch submission reduce consumption without changing what an engineer can attempt, so a slowdown usually points to a routing rule sending hard work to a small model. Measure the value of merged changes on both sides of the change and the answer stops being a matter of opinion.

