Next Steps
Module 02 · Context Engineering /Lesson 2.1 /6 min

The context budget as a design artefact

Decide what goes in the window before the session starts, the same way you would decide an interface before writing the class.

Most people assemble context reactively: start a session, let the agent read whatever it decides to read, notice halfway through that half of it was irrelevant. The alternative is to treat the window as a designed artefact with a budget, filled deliberately.

This is not ceremony. It takes ninety seconds and it is the highest-return ninety seconds available in agent work.

The four slots

Every productive session's context divides into four kinds of material. Naming them lets you notice when one is missing and when one has swallowed the others.

SlotContainsTypical budget
StandingSystem prompt, project brief, tool definitions3–12k, fixed, paid every session
ReferenceThe exemplar, the interface, the convention you must match1–5k
WorkingThe files actually being changed, and their tests3–15k
TaskIntent, scope, constraints, definition of done200–600 tokens

Look at the last row. The task slot is the smallest by an order of magnitude and it determines almost everything about the outcome. Most sessions that go wrong have a large working slot and a task slot of one sentence.

The budget line

Before a session, write the budget. Literally, in the scratch file:

notes/current-task.md
TASK    Make the CSV importer reject rows with a malformed date rather
        than silently dropping them.
SCOPE   src/import/parse-row.ts, src/import/__tests__/parse-row.test.ts
REF     src/import/parse-header.ts  (this is how we report row errors)
WORK    the two scope files
NOT     the schema, the error taxonomy, the CLI output format
DONE    npm run check passes; a test asserts the error for 01/13/2026

BUDGET  ~7k loaded before the plan step. If I am past 15k, I have loaded
        something I do not need.

The budget number is the useful part. It converts "am I loading too much?" from a judgement call made under pressure into a check you can run.

What earns a place

Three questions, applied to every candidate:

  1. Will the agent change it? Then it is working context and it goes in whole.
  2. Will the agent need to match it exactly? Then it is reference and goes in whole. An exemplar half-loaded is worse than none: the agent fills the gap from its prior.
  3. Does the agent only need to know it exists? Then it is a pointer, not content. "the retry cap lives in config/queue.ts" is twelve tokens and does the job.

Anything that fails all three is not context. It is ballast that will cost you attention for the rest of the session.

The exploration exception

Sometimes you do not know what the working set is. That is a different kind of session with a different budget, and the crucial rule is that it is a different session.

two sessions, not one
session 1  EXPLORE     budget: generous, 40-60k is fine
                       output: a written answer, 15 lines, in a file
                       then: CLEAR

session 2  IMPLEMENT   budget: tight, the 15 lines plus the working set
                       output: a diff

Doing both in one session means implementing from the far end of a 60k window, with all the exploration debris still competing for attention. The written answer is the whole point: it is a dense, high-signal summary that costs 400 tokens instead of the 50k it took to produce.

Reading your own budget after the fact

At the end of a session, ask one question: what is in this window that I would not load again? Every answer is a habit to fix. Common ones:

  • A file read to check one thing, never referenced again — should have been a grep.
  • Full test output from a run that passed — should have been quiet.
  • A directory listing — should have been a targeted search.
  • Three turns of a debugging detour, resolved — should have been its own session.

This audit takes thirty seconds and the same three or four answers will recur for a fortnight before they stop.

Exercise

Write the budget block above for your next three real tasks, before opening an agent. Include the budget number.

Afterwards, record the actual tokens loaded before the plan step against your estimate. The gap between the two is the interesting number — most people are out by 3–5x on their first attempt, and consistently within 30% by the tenth.

Worked solution

Three real tasks, estimated and measured.

task 1 - "add a rate limit to the login endpoint"
estimated  6k
actual     31k
gap        the agent read the whole middleware directory (7 files) looking
           for "where rate limiting would go", because SCOPE named a file
           that did not exist yet.

fix        when the change creates a new file, REF must name the closest
           existing sibling. Added:
             REF  src/middleware/csrf.ts (same shape, same registration)
           re-run: 8k actual.
task 2 - "fix the duplicate-email crash on signup"
estimated  5k
actual     22k
gap        no reproduction in the task block, so the agent explored to
           find one - four files and a full test run.

fix        DONE lines are cheap; reproductions are cheaper than exploration.
             REPRO  npm test -- signup.test.ts -t "duplicate"  (currently fails)
           re-run: 6k actual.
task 3 - "extract the PDF renderer into its own module"
estimated  12k
actual     14k
gap        small, and in the right direction.

The pattern in the two misses is the same and it is worth stating: the agent explores when the task block leaves it a question. Task 1 left "where does this go"; task 2 left "how do I see the bug". Every unanswered question in the task slot gets paid for out of the working slot, at roughly twenty times the price.

So the practical refinement to the budget block: before you start, read it back and ask what a competent stranger would still have to go and find out. Answer that in the block, in a line, and the budget holds.

Takeaways

  • Context divides into standing, reference, working and task — the smallest slot decides the most.
  • Write a token budget before the session; it turns a judgement call into a check.
  • Explore and implement in different sessions, carrying a written summary rather than the exploration.
  • Every question your task block leaves open gets paid for in exploration, at roughly twenty times the price.

Check yourself

Your loaded context came out at five times your budget. What is the most likely cause?

A course by