Next Steps
Reference/97 terms

Glossary

The vocabulary both courses use. Free to read, and worth reading even if you never buy anything — most bad agent sessions come from confusing two of these terms with each other.

The model

Model
The parameters. Stateless: it performs next-token prediction and nothing else. Everything that resembles memory or agency is supplied by the layer around it.
Forward pass
One processing of the entire context to produce a probability distribution over the next token. A 500-token answer is 500 passes over your whole context.
Inference
Running a trained model to produce output — what happens on every request to a provider.
Token
The atomic unit a model reads and writes. Roughly four characters of English, fewer for code, far fewer for minified or encoded data.
Tokenisation
Splitting text into tokens. Trained mostly on natural language, which is why code costs roughly 30% more tokens per character than prose.
Effort
A dial controlling how much internal reasoning happens before the visible answer. High for deciding, low for doing. It amplifies context, including its noise.
Sampling
Choosing one token from the distribution a forward pass produces. Governed by temperature and top-p where those are exposed.
Non-determinism
The same input can produce different output. Comes from sampling and from provider-side batching, hardware and routing — temperature zero is not determinism.
Input tokens
What you send on each request. Cheap per token, enormous in volume, and re-sent in full every turn.
Output tokens
What the model generates. Priced higher per token and usually under 5% of a session's volume.
Prefix cache
Provider-side storage of the processed, unchanged start of your context, billed at a large discount. It matches from position zero and stops at the first difference.
Cache tokens
Input tokens served from the prefix cache. A healthy long session is 90%+ cached on later turns.

Sessions and context

Harness
The program around the model: system prompt, tools, permission model and context assembly. Most comparisons between "AI coding tools" are comparisons of harnesses.
Agent
A model harnessed with tools, a system prompt and a context window, taking turns with a user.
Context window
Everything the model sees on one request. Finite, model-specific, and re-sent in full every time.
Context
The relevant information the agent has right now. Not the same as the window: a full window can hold terrible context.
Session
One bounded run of interaction. Starts empty, accumulates, ends when cleared. The real unit of agent work.
Turn
One user message plus everything the agent does before handing control back — possibly dozens of tool calls.
System prompt
Instructions the harness prepends to every request, including every connected tool's definition.
Stateless
Carries nothing forward between requests. The model is stateless; apparent memory is the harness re-sending history.
Session shape
Which kind of work a session is doing: explore, decide, implement, verify or repair. Each needs different context, budget, permissions and ending.
Loaded-before-decision
Tokens in the window at the moment the agent proposes its plan. The single best predictor of session quality.
Signal ratio
The fraction of the window that still bears on the work. Under 30% means the session did more than one job.

Tools and environment

Tool
A function the harness exposes for the agent to call: read, write, search, run a command.
Tool call
Structured text the model generates naming a tool and its arguments. The harness recognises and executes it; there is no special mechanism.
Tool result
What the harness appends after executing a call. It stays in the window permanently and is re-sent every subsequent turn.
MCP
A protocol for plugging external tool servers into a harness. Each connected server's definitions sit in every session's system prompt.
Permission mode
Which tool calls run without asking. Pre-approve the frequent and safe, so the remaining prompts get read.
Approval fatigue
Approving prompts reflexively because there are too many. A control exercised without being read is not a control.
Sandbox
An isolated environment the agent runs inside — container, VM, worktree, restricted shell. Lowers blast radius so autonomy can be raised.
Blast radius
What a mistake could reach. A separate dial from autonomy, and the one to lower first.
Hook
Code that runs at a fixed point in the agent loop — before a tool call, after an edit, before a commit. Enforcement that costs no context and does not degrade.
Worktree
A second checkout of the same repository on another branch. Near-free isolation that makes parallel agent work safe.

Failure modes

Hallucination
Confidently-wrong output. Two flavours: factuality (invented from training) and faithfulness (drifted from loaded context).
Parametric knowledge
What the model absorbed during training, frozen in its parameters. The source of factuality errors.
Knowledge cutoff
The date past which the model has no parametric knowledge. Anything that changed since is described confidently in its previous state.
Contextual knowledge
Facts read from the current context rather than recalled. Always prefer it for anything specific to your system.
Attention budget
Each token has a fixed amount of influence to distribute across the rest of the context. The weights sum to one, so context is zero-sum.
Attention degradation
Signal on the relationships that matter weakens as the window fills with relationships that do not. Gradual, silent, and with no error message.
Smart zone
The early part of a session where the agent is sharp, recalls instructions and stays in scope. Budget the zone, not the window.
Dumb zone
What follows: forgotten constraints, repeated mistakes, drift toward generic idiom. Same model, same prompt — only context volume differs.
Context poisoning
Material in the window that actively steers wrong rather than merely diluting: an abandoned approach, a corrected mistake, a stale file, an adjacent fact, an injected instruction.
Drift
Output pulled toward the training distribution — generic idiom rather than your conventions — as your codebase's share of the window falls. Self-reinforcing once drifted code lands.
Sycophancy
Agreeable output produced because agreement was rewarded in training. Ask for the case against, not for confirmation.
Speculative generality
Extension points, configuration and abstractions for cases that do not exist. The characteristic failure of delegated design.
Prompt injection
Instructions inside content the agent reads — an issue, a page, a CSV cell — being followed. Possible because there is no privileged channel.

Handoffs

Clearing
Ending a session and starting with an empty context. The main hygiene practice, not an admission of failure.
Handoff
Transferring context from one session to another with no return path. Everything that matters must be in the artefact.
Handoff artefact
The document that carries context across the gap: a spec, a ticket, a decisions file or a handoff note.
Spec
A handoff artefact describing multi-session work: goal, exclusions, constraints, decisions with reasons, acceptance criteria. Free of session-specific detail.
Ticket
A handoff artefact scoping exactly one session, executable by an agent with no memory of any previous one.
Compaction
Summarising earlier history to free room. Converts primary sources into secondary ones and drops reasons and negative results, silently.
Autocompact
Compaction triggered automatically as the window fills. An alarm that you are long past the smart zone, not a feature.
Primary source
The thing itself: the actual file, the real schema, the raw output. Complete and current, expensive to load.
Secondary source
An account of a primary: a README, a summary, a design doc, anything the agent summarised earlier. Cheap, lossy, and believed.
Decisions file
An append-only record of decisions with their reasons and dates. The reason decays faster than the decision, which is why they are kept together.

Memory and steering

AGENTS.md
The project's standing brief, loaded at session start. Names differ by harness; the role does not. Belongs in version control.
Progressive disclosure
Loading only what is needed now, with pointers to the rest. Applies to documentation, data and agent output alike.
Context pointer
A mention telling the agent where to look and under what condition. A pointer without a trigger is read always or never.
Skill
A named procedure kept out of context until a pointer activates it. For work that applies to some sessions, not most.
Exemplar
An existing file or diff used as the specification for new work. Communicates the conventions nobody has written down.
Subagent
An agent spawned by another, running in its own window. Buys a clean context; costs everything the parent knows.
Orchestrator
An agent holding a plan and dispatching work to others. Usually better as a shell loop with a plan file.
The check ladder
Four rungs of enforcement: said in the session, written in the brief, enforced by a check, impossible by construction. The skill is moving rules up.

Verification

Automated check
Deterministic verification: types, lint, tests. Costs no context, never tires, runs inside the agent's own correction loop.
Check command
The single fast command the agent runs to find out whether it broke something. Speed determines frequency, and frequency is the whole mechanism.
Automated review
A fresh session judging a diff it did not write. Non-deterministic, excellent at mechanical failure modes. A filter, not a gate.
Action rate
The fraction of an advisory channel's findings that anyone acts on. Below 40% it will be ignored within a month.
Adversarial verification
A second pass whose job is to refute each finding. Roughly halves false positives, which is what keeps people reading.
Characterisation test
A test capturing what code currently does, including what looks wrong. The only way to build a net around behaviour nobody can state.
Property test
A test asserting something holds for generated inputs. Covers the cases nobody thought of, which is exactly the gap generated code leaves.
Mutation testing
Deliberately changing code to see whether any test fails. Tells you what coverage cannot: which behaviours no test distinguishes.
Differential testing
Running old and new implementations against the same inputs and comparing. The strongest check available for a refactor or migration.
Golden file
A committed known-good output, diffed on every run. Catches serialisation-level changes that schema checks miss.
Contract snapshot
A generated, committed record of a published surface — API schema, exported types, event payloads. Turns an invisible break into a diff someone must approve.
Shadow mode
Running a new implementation alongside the old, reporting divergences, with the old one authoritative. What makes strangling a legacy system safe.
Defect escape rate
Defects reaching production per week or per feature. The one metric that cannot be improved by working faster.
Review depth
The proportion of changed lines somebody genuinely read. Falls silently under volume; nobody reports it.
Trust calibration
Matching review depth to check coverage and blast radius, per area. Drifts toward over-trust because failures are rare and delayed.
Audit sample
Properly reviewing something you approved with a skim. The only instrument that measures your own judgement rather than the system's.

Patterns of work

Human-in-the-loop
Pairing: watching tool calls, correcting plans, reviewing as it goes. Buys correction, costs attention.
AFK
Initiating and walking away. Buys throughput, requires a complete spec, stop conditions and a contained environment.
Vibe coding
Accepting output without reading it. Fine for throwaway work; expensive the moment the code has a second reader.
Grilling
Having the agent interview you, one question at a time, until the requirement is pinned down. The highest-leverage specification technique.
Prototyping
Building a rough version to answer one named question, then deleting it. Now cheap enough to beat another round of discussion.
Stop condition
A stated situation in which halting and reporting is the correct outcome. Without one, a blocked unattended agent works around the obstacle.
The four evasions
Widening a type, deleting an assertion, skipping a test, adding a suppression — the moves an agent makes when the honest path to a green check is blocked.
Poison probe
Asking the agent to list every rejected approach and every instruction still in force. Diagnoses a compromised window in one turn.
Canary constraint
A distinctive, checkable instruction used to measure where your own dumb zone starts.
Eval suite
A small fixed set of tasks from your own history, run against a model or a setup change. An afternoon once, an answer in an hour thereafter.

People and systems

DX
Developer experience: how well a codebase lets humans work.
AX
Agent experience: how well an environment lets agents perform. Mostly overlaps with DX, with searchability, locality and explicitness weighted higher.
Hyrum's law
With enough consumers, every observable behaviour of your system is depended on by someone, regardless of what you documented.
Strangler pattern
Replacing a system incrementally behind a facade, routing one piece at a time, deleting the old piece once the new one has proven itself.
Seam
A place where behaviour can be substituted without editing the surrounding code. Finding one is most of the work in testing legacy code.
ADR
Architecture decision record: a dated record of what was decided and why. The one prose format that does not go stale.
Bus-factor question
"What would we be stuck on if one person were unavailable for two weeks?" The answers are your runbook backlog, in priority order.