Where the money actually goes
A cost model precise enough to act on, and the three line items that dominate every bill.
Most people's intuition about agent cost is wrong in the same direction: they assume the expensive part is the code being written. It is not. Output tokens are a small fraction of almost every bill.
The arithmetic
Because the whole context is re-sent on every request (module 01), a session's cost is the sum over every turn of everything in the window at that point.
turn window input (uncached) output cumulative input
1 6k 6,000 400 6,000
5 38k 4,200* 600 62,000
10 71k 3,100* 500 118,000
20 140k 5,400* 700 294,000
--------
* most input is cached after turn 2
total input 294,000 (of which ~87% cached)
total output 11,000Two observations. Output is under 4% of the token volume. And without prefix caching this session would have cost roughly eight times more in input — caching is what makes long sessions affordable, and it is also what removes the financial pressure that used to keep them short.
The three dominant line items
- Tool results that stay in the window. One verbose test run at turn 3 is re-sent on every subsequent turn. An 8,000-token output on turn 3 of a 20-turn session costs roughly 136,000 input tokens over the session.
- Files read and never used. Same mechanism. The four files read during exploration that turned out to be irrelevant are paid for on every remaining request.
- Always-loaded overhead. Tool definitions and instruction files, on every session. Five chatty MCP servers at 20,000 tokens is 20,000 tokens times every session you run, forever.
All three are context-management problems, which is the point: the cost lever and the quality lever are the same lever. Everything in module 02 reduces both.
What does not move the bill much
- Prompt length. A 400-token opening message is noise against a 300,000-token session.
- Output length. Asking for shorter answers saves single-digit percentages.
- Effort, on most turns. Worth using deliberately (module 01) but not a primary cost driver unless it is on by default.
The one-hour audit
1. Start a fresh session. Record the token count before typing.
-> anything above ~10k is always-loaded overhead worth attacking.
2. Run your five most common commands. Record each one's output size.
-> anything over 500 tokens gets a quieter invocation.
3. Take a typical session and list what is in the window at the end.
-> the fraction that is still relevant is your efficiency.
4. Multiply the waste by sessions per week.A worked example of the multiplier
npm test prints 6,200 tokens instead of 60. It runs ~4 times per session, at an average of turn 6 of 14. per session: 6,140 extra x 4 runs x ~8 remaining turns = ~196,000 tokens per week: x 15 sessions = ~2.9M per year: = ~135M
One line of configuration. This is the shape of nearly every real saving: not a change in how you work, but a change in what your tools return.
Budgeting without micromanaging
Per-session cost is not worth tracking. Two numbers are:
- Cost per feature shipped, monthly. It should fall as your process improves; a rise means sessions are getting longer or specs thinner.
- The outlier sessions. The one that cost five times the median is always diagnosable, and the cause is almost always a directory dump, a verbose command, or an uncleaned detour.
Chasing individual session costs is a bad use of attention. Finding the systematic leak once and fixing it in configuration is a good one.
Exercise
Run the four-step audit. Record your fresh-session baseline, the output size of your five most common commands, and the end-of-session relevance fraction.
Fix the worst tool output, then compute the annual multiplier for that one change. Most people find a seven-figure token saving from a single line of configuration.
Worked solution
1. fresh session baseline: 23,400 tokens
system prompt + harness ~3,100
AGENTS.md ~2,000
MCP tool definitions ~18,300 <-- three servers,
one of them used
2. command output sizes
npm test 6,200 <--
npm run build 3,400
git status (dirty repo) 800
npm run check 140
rg 180
3. end-of-session relevance: 62k window, ~14k still relevant = 23%
4. sessions per week: 18 a) disconnected two unused MCP servers
baseline 23,400 -> 6,900
saving: 16,500 x 18 sessions x 46 weeks = ~13.7M tokens/year
and the fresh session now starts with 16k more attention budget
b) npm test -> --reporter=dot --bail=1 2>&1 | tail -40
6,200 -> 210
ran ~4x per session at average turn 6 of 14
saving: ~5,990 x 4 x 8 x 18 x 46 = ~127M tokens/yearAbout 140 million input tokens a year, from two configuration changes and twenty minutes of work. At typical cached-input pricing that is a meaningful but not life-changing amount of money — and that is the wrong way to read it.
The number that mattered was the third one: end-of-session relevance at 23%. Three quarters of every window was material that had stopped bearing on the work. After the fixes, plus the clearing discipline from module 02, that went to 58% over the following month. The bill halved as a side effect; the output got noticeably better, which was the actual point.
The one thing I would do differently: I spent the first hour looking at per-session costs in the usage dashboard, which told me nothing. The audit above takes twenty minutes and finds the systematic leaks, which is where all the value is.
Takeaways
- Output tokens are under 5% of a typical bill; context re-sent every turn is the cost.
- A verbose tool result is charged on every remaining turn of the session, not once.
- Always-loaded overhead — tool definitions, instruction files — is charged on every session you ever run.
- The cost lever and the quality lever are the same lever.
Check yourself
Why does one verbose test run early in a session cost so much more than its own size?
The context is stateless and re-sent in full each turn. A 6,000-token result at turn 3 of a 14-turn session is paid for eleven more times, and it competes for attention on every one of them.
A course by Pieter Zandbergen