Summary
The price of a million tokens has fallen steeply since 2023. Yet teams building software with AI are spending more each quarter, not less. The reason is volume: agentic coding tools plan, read whole codebases, run tests and retry, so a single task can consume many times the tokens of a chat exchange.
Most cost tooling was built for the AI inside a shipped product. It meters customer traffic through a gateway. The tokens a team burns while building that product sit elsewhere, spread across seats, tools and personal API keys, and are visible only as a monthly total.
The paper argues that build-time AI spend should be managed like any other project cost: budgeted at kickoff, tracked as burn-down, capped at a ceiling and charged back to the client or cost centre that incurred it.
Inside the paper
- How agentic tools multiply tokens per task
- Why runtime gateways don't cover build-time spend
- What a per-project budget, burn-down and ceiling look like in practice