Agent Engineering
Module 02 · Concepts/Lesson 2.2/3 min

Tokens, windows, and the bill

A working numerical intuition for what a session costs and why the second half costs more than the first.

A token is the atomic unit a model reads and writes — roughly three-quarters of an English word, less for code, much less for minified JSON or base64. Useful rules of thumb: 1 token ≈ 4 characters of English, a 400-line source file ≈ 4–6k tokens, a dense stack trace ≈ 1–3k.

The context window is a per-request budget

The context window is everything the model sees on one request: system prompt, standing instruction files, the entire conversation so far, every tool result, and your latest message. It is finite and model-specific — commonly a few hundred thousand tokens on current frontier models.

The important and counter-intuitive part: the whole window is re-sent on every single request. The model is stateless, so there is no "continue from where we were". Turn twelve sends turns one through eleven again, in full.

What that does to cost

a session's token arithmetic
turn 1   sends  6k   (system prompt + brief + your message)
turn 5   sends  38k  (…everything above, plus four turns of tool results)
turn 20  sends  140k (…all of it again)

total input tokens for a 20-turn session is not 140k.
It is the SUM of every request: often 1.2-1.8M.

So a long session is quadratic-ish in cost, not linear. This is the economic argument for short sessions, and it happens to point the same direction as the quality argument in Lesson 2.4.

Input, output, and cache

  • Input tokens — what you send. Cheap per token, enormous in volume.
  • Output tokens — what the model generates. Several times more expensive per token, but far fewer of them.
  • Cache tokens — providers keep a prefix cache of the unchanged start of your context, and bill the cached part at a large discount. This is why the stable parts of your context (system prompt, project brief) are nearly free after the first request, and why changing something early in the context is expensive: it invalidates the cache for everything after it.

Watch out

The prefix cache has a practical consequence for how you work: appending to the conversation is cheap, but editing an instruction file mid-session busts the cache and re-bills the whole prefix. If you need to change standing instructions, change them and start a fresh session — you were going to want that anyway.

Reading your own usage

Most harnesses can show per-session token counts. Find that display and glance at it for the next week. The goal is a physical intuition for what a big file read costs, so that "let me just cat the whole schema" stops feeling free.

Try it

Start a fresh session. Note the token count. Ask the agent to read one medium-sized source file. Note it again. Then ask it to run your test suite and note it a third time. Write the three numbers in your log — particularly the test-suite delta, which surprises most people.

Takeaways

  • The entire context window is re-sent on every request; long sessions cost far more than their final size suggests.
  • Output tokens are priced higher per token, but input volume usually dominates the bill.
  • The prefix cache makes stable context nearly free and mid-session instruction edits expensive.
Why does asking an agent to run a verbose test suite twice cost more than the two runs suggest?

Because each run’s full output is appended to the history and then re-sent on every subsequent request for the rest of the session. A 3k-token test output is not a one-off 3k charge; it is 3k added to the per-request cost of every remaining turn. Prefer commands with quiet output, and prefer failing fast.

A course by Pieter Zandbergen