Cut the token bill. Keep the output.
Chameleon routes every task to the cheapest model that can do it, and what you spend is priced against graded output. The budget conversation ends with evidence.
Trusted by engineering teams




1,284 tasks routed · 2 held at frontier
- Haiku 4.5Mechanical, fully specified$0.02−94%
Scaffold CRUD endpoints
navigaracom/vision · ENG-2917
- Sonnet 5Touches security, needs care$0.28−61%
Refactor auth middleware
navigaracom/vision · ENG-2884
- Opus 5High blast radius, no downgrade$1.42
Payments architecture spike
navigaracom/billing · ENG-2902
- Haiku 4.5Pattern repetition$0.03−91%
Backfill unit tests
navigaracom/identity · ENG-2871
- Opus 5Ambiguous, spans four services$1.18
Incident root cause, checkout
navigaracom/vision · INC-441
Chameleon picks the model. You teach it what best means.
Most work does not need your most expensive model. Some must never run on anything else. Chameleon reads each task and its blast radius, then routes it to the cheapest model that holds the quality bar.
Scaffold CRUD endpoints
navigaracom/vision · ENG-2917
- Haiku 4.592$0.02
- Sonnet 597$0.11
- Opus 599$0.98
Haiku 4.5 at $0.02. 98% cheaper than defaulting to the frontier model.
Per-task, not per-seat
The unit of routing is a task with a known shape, not a developer with a licence.
Trainable, not fixed
Every accept, revert and rollback is a label. Your policy diverges from our default on purpose.
Refusal is a feature
Work with real blast radius stays on the frontier model, and the router records why it declined to save money.
The default policy is ours. The one you run is yours.
Every codebase has its own idea of which tasks are dangerous, so the policy is trainable rather than a vendor default.
Decide migrations always run on the frontier model and that becomes policy. The router argues with the bill, never with you.
A cheaper model on work nobody asked for is cheaper waste.
Two ways to think about AI efficiency: which models you run, and what you spend the tokens on. Chameleon handles the first, and it caps out. Model choice saves a percentage. Work type decides whether the spend should exist.
A pull request is not a unit of value.
Routing cuts the bill. Whether it cut the output depends on what you divide by. Spend per pull request measures activity, and activity is what inflates when work gets easier to produce. Divide by it and every rollout looks like a win.
- ActivityPull requests mergedWhat the rest of the category divides byDenominator74210 PRs+184% growth$1,688$595per pull requestOverstates the gain
- Graded valueETV, scored per commitA language model reads each diff and grades what it was worthDenominator46.675.4 ETV+62% growth$2,680$1,656per ETVThe honest number
Activity inflates when the work gets cheaper to produce. Graded value does not, so it is the only denominator that survives an AI rollout.
ETV is a per-commit value score: a language model reads each diff and grades it. Scored, not counted.
Stop defending the bill. Start defending the return.
A budget increase is easy when the price of a unit of shipped work falls while volume rises. Spend breaks down the same way output does, tied to the objectives it moved.
Cost per ETV down 38%
Absolute spend rose over the same period. What fell is the price of a unit of graded output, which is the number that settles the argument.
- Features$9,48552%Shipped against a roadmap objective
- Maintenance$4,01322%Keeping what exists running
- Tests$2,18912%Coverage written alongside the change
- Docs$9125%Runbooks and API reference
- Fixes$1,6419%Bugs, regressions, edge cases
At today’s rate, another $40,000 a quarter buys roughly 165 ETV of additional shipped work.
Follow a dollar from the initiative to the commit.
Spend flows from an initiative down to a commit, and at every link it either still answers to the mandate above it or it does not. The band at the bottom left is the 21% with no issue behind it.
Green clears the mandate above it. Amber is questionable. Red has no mandate behind it. Band width is cost.
Board-ready exports, an audit trail from every dollar to the commits behind it, and per-team breakdowns finance can reconcile.
Knowing what you spent is not the same as knowing what to fix.
The usual answer is one maturity score that tells nobody what to do on Monday. We score five dimensions from evidence, compare each against your org median, and name the one gap worth closing next.
Five independent levels for how this team works with AI agents. Each level comes from its own evidence, so the levels are never added together.
Ahead of the median on 3 of 5 dimensions. Largest lead: Delegation depth (+1.8). Largest gap: Context leverage (−1.5).
- Agent fluency4.4AI-active 7 of 8 weeks
- Delegation depth3.9Deep delegation on 21 of 28 agent days
- Roadmap alignment3.572% of work credited to roadmap objectives
- Context leverage0.83 of 12 artifacts state a why and outcome
- Agentic autonomy0.52% of work owned end-to-end by agents
- Require a why and an expected outcome on every agent task, not just a title
- Attach the ticket and the failing test to the prompt so agents stop rediscovering context
- Promote the three artifacts that already do this into templates for the rest of the team
The levels are never added together. A team that delegates deeply but writes no context has a specific, fixable problem, and an average hides it.
What this changes.
Answer the CFO’s question about the AI line item with a cost per unit of shipped work, not a per-seat count
Walk into the budget review knowing what another $40,000 a quarter actually buys
Show which teams turn tokens into roadmap work and which turn them into waste
Give a team one named thing to improve next quarter, with the evidence behind it
See what your tokens
actually bought.
Measurement runs read-only against your repositories, boards and provider bills. Routing is the one place Navigara sits in the request path, and it is opt-in, per team.
Your keys, your providers
Chameleon routes through your own provider accounts. No resale, no markup.
Nothing retained
Prompts and completions are not stored once a routing decision is made.
Every decision logged
The model chosen, the reason, the cost, and the cheaper option it rejected.