The Token Paradox

Why AI gets cheaper per token and more expensive per project.

Summary

The price of a million tokens has fallen steeply since 2023. Yet teams building software with AI are spending more each quarter, not less. The reason is volume: agentic coding tools plan, read whole codebases, run tests and retry, so a single task can consume many times the tokens of a chat exchange.

Most cost tooling was built for the AI inside a shipped product. It meters customer traffic through a gateway. The tokens a team burns while building that product sit elsewhere, spread across seats, tools and personal API keys, and are visible only as a monthly total.

The paper argues that build-time AI spend should be managed like any other project cost: budgeted at kickoff, tracked as burn-down, capped at a ceiling and charged back to the client or cost centre that incurred it.

Inside the paper

  • How agentic tools multiply tokens per task
  • Why runtime gateways don't cover build-time spend
  • What a per-project budget, burn-down and ceiling look like in practice