Next Steps
Module 01 · Mechanics You Can Reason About /Lesson 1.1 /6 min

What a forward pass actually does

Enough mechanism to stop guessing. Not a maths lesson — a working model of the thing you are delegating to.

You do not need the linear algebra. You do need a model of the machine accurate enough that its behaviour stops being surprising, because almost every technique in this course is a consequence of one of four facts about how a request is processed.

One request, one pass

When your harness sends a request, it sends a single flat sequence of tokens: system prompt, then your instruction files, then every turn of the conversation, then every tool result, then the newest message. The model has no notion that these came from different places or at different times. To it, there is one document, and the job is to continue it.

That sequence is processed once to produce one token. Then the whole sequence plus that token is processed again to produce the next. And again. A five-hundred-token answer is five hundred passes over your entire context.

the loop, honestly
context = [system][brief][turn 1][tool result][turn 2]...[your message]
          |
          v
     forward pass  ->  probability distribution over the next token
          |
       sample one  ->  append to context
          |
       repeat until an end-of-turn token is sampled

First consequence: everything is input. There is no separate channel for instructions, no privileged region. The rule you wrote in AGENTS.md and the stack trace a tool returned are the same kind of object competing on the same terms. This is why a large tool result can drown a standing instruction, and why prompt injection works at all.

Every token sees every other token

Inside the pass, each token computes a weighted view of every other token — attention. The weights are learned, not fixed: a variable name attends strongly to its declaration, a closing brace to its opening one.

The weights sum to one. That is the whole of the next lesson, but state it now: attention is a fixed budget that gets divided, not a resource that grows with the context. Adding irrelevant tokens does not leave the relevant relationships untouched; it makes each of them a smaller share of a fixed total.

Second consequence: "just include it, it might help" is not free and not neutral. It has a cost paid by everything else in the window.

The model is a distribution, not an answer

A forward pass does not produce a token. It produces a probability over every token in the vocabulary — typically a hundred thousand or more of them. Something downstream picks one.

So when an agent writes import { retry } from './utils' and no such export exists, nothing went wrong mechanically. Given your codebase, that continuation was highly probable. The machine did exactly what it does. It has no separate faculty that checks whether a plausible continuation is a true one.

Third consequence: correctness is not a property the model provides. It is a property your environment has to impose, which is why so much of this course is about checks.

Nothing carries over

After the pass, no state is kept. The next request rebuilds everything from the text it is sent. "Remember what I said earlier" works only because the harness re-sends what you said earlier.

Fourth consequence: every apparent memory is a design decision someone made about what to include. Memory systems, compaction, instruction files, handoff notes — all of them are answers to the same question: what goes back in the box next time. That question is yours to answer, and it is the highest-leverage one you have.

The four facts, together

FactWhat follows
Everything is one undifferentiated sequenceInstructions compete with content; volume dilutes rules
Attention is a fixed budgetIrrelevant context actively degrades relevant reasoning
Output is a distribution, sampledPlausibility, not truth; verification must be external
No state survives a requestContext assembly is the entire lever you control

Hold on to these. Every later module is an application of one of them, and when you hit a behaviour this course does not cover, working back to which of the four is in play will usually tell you what to do.

Exercise

Take the last agent session that went badly for you. Write one sentence classifying it under exactly one of the four facts — instruction drowned by volume, attention diluted, plausible-but-false output, or context that was never re-sent.

Then write the intervention that fact implies. Not "be more careful": the specific change to what goes into the window, or the specific check that would have caught it.

Worked solution

Worked example. The failure: "I told it not to use the legacy client, and forty minutes later it used the legacy client."

Classification. Tempting to call this fact three — the model was wrong. It is not. The instruction was given once, early, in a window that then filled with forty minutes of file reads. By the time the decision was made, that one sentence was a vanishing share of the attention budget. This is fact two, with fact one making it worse: the instruction was competing on equal terms with thousands of tokens of code that all demonstrated the legacy client being used.

Implied intervention. Not a longer instruction — that adds volume, which is the direction that caused the problem. Two real fixes, in order of strength:

  1. A lint rule banning imports of the legacy client outside its own directory. Deterministic, costs no attention, and fails inside the agent's own check loop so it self-corrects. This is the fact-three answer applied to a fact-two problem: impose correctness externally.
  2. Failing that, the instruction moves into the always-loaded brief and the session gets split so the decision happens in the first ten minutes rather than the fortieth.

The general shape: when a rule is being forgotten, stop trying to say it better and start trying to say it somewhere that does not depend on attention.

Takeaways

  • One request is one flat sequence; your instructions have no privileged status within it.
  • Attention is a fixed budget divided among all tokens, so irrelevant context costs the relevant context directly.
  • The model emits plausibility; correctness has to come from your environment.
  • Nothing persists between requests — context assembly is the whole lever.

Check yourself

An agent ignores a constraint you stated clearly at the start of a long session. Which fact best explains it?

A course by