Reference/97 terms
Glossary
The vocabulary both courses use. Free to read, and worth reading even if you never buy anything — most bad agent sessions come from confusing two of these terms with each other.
The model
- Model
- The parameters. Stateless: it performs next-token prediction and nothing else. Everything that resembles memory or agency is supplied by the layer around it.
- Forward pass
- One processing of the entire context to produce a probability distribution over the next token. A 500-token answer is 500 passes over your whole context.
- Inference
- Running a trained model to produce output — what happens on every request to a provider.
- Token
- The atomic unit a model reads and writes. Roughly four characters of English, fewer for code, far fewer for minified or encoded data.
- Tokenisation
- Splitting text into tokens. Trained mostly on natural language, which is why code costs roughly 30% more tokens per character than prose.
- Effort
- A dial controlling how much internal reasoning happens before the visible answer. High for deciding, low for doing. It amplifies context, including its noise.
- Sampling
- Choosing one token from the distribution a forward pass produces. Governed by temperature and top-p where those are exposed.
- Non-determinism
- The same input can produce different output. Comes from sampling and from provider-side batching, hardware and routing — temperature zero is not determinism.
- Input tokens
- What you send on each request. Cheap per token, enormous in volume, and re-sent in full every turn.
- Output tokens
- What the model generates. Priced higher per token and usually under 5% of a session's volume.
- Prefix cache
- Provider-side storage of the processed, unchanged start of your context, billed at a large discount. It matches from position zero and stops at the first difference.
- Cache tokens
- Input tokens served from the prefix cache. A healthy long session is 90%+ cached on later turns.
Sessions and context
- Harness
- The program around the model: system prompt, tools, permission model and context assembly. Most comparisons between "AI coding tools" are comparisons of harnesses.
- Agent
- A model harnessed with tools, a system prompt and a context window, taking turns with a user.
- Context window
- Everything the model sees on one request. Finite, model-specific, and re-sent in full every time.
- Context
- The relevant information the agent has right now. Not the same as the window: a full window can hold terrible context.
- Session
- One bounded run of interaction. Starts empty, accumulates, ends when cleared. The real unit of agent work.
- Turn
- One user message plus everything the agent does before handing control back — possibly dozens of tool calls.
- System prompt
- Instructions the harness prepends to every request, including every connected tool's definition.
- Stateless
- Carries nothing forward between requests. The model is stateless; apparent memory is the harness re-sending history.
- Session shape
- Which kind of work a session is doing: explore, decide, implement, verify or repair. Each needs different context, budget, permissions and ending.
- Loaded-before-decision
- Tokens in the window at the moment the agent proposes its plan. The single best predictor of session quality.
- Signal ratio
- The fraction of the window that still bears on the work. Under 30% means the session did more than one job.
Tools and environment
- Tool
- A function the harness exposes for the agent to call: read, write, search, run a command.
- Tool call
- Structured text the model generates naming a tool and its arguments. The harness recognises and executes it; there is no special mechanism.
- Tool result
- What the harness appends after executing a call. It stays in the window permanently and is re-sent every subsequent turn.
- MCP
- A protocol for plugging external tool servers into a harness. Each connected server's definitions sit in every session's system prompt.
- Permission mode
- Which tool calls run without asking. Pre-approve the frequent and safe, so the remaining prompts get read.
- Approval fatigue
- Approving prompts reflexively because there are too many. A control exercised without being read is not a control.
- Sandbox
- An isolated environment the agent runs inside — container, VM, worktree, restricted shell. Lowers blast radius so autonomy can be raised.
- Blast radius
- What a mistake could reach. A separate dial from autonomy, and the one to lower first.
- Hook
- Code that runs at a fixed point in the agent loop — before a tool call, after an edit, before a commit. Enforcement that costs no context and does not degrade.
- Worktree
- A second checkout of the same repository on another branch. Near-free isolation that makes parallel agent work safe.
Failure modes
- Hallucination
- Confidently-wrong output. Two flavours: factuality (invented from training) and faithfulness (drifted from loaded context).
- Parametric knowledge
- What the model absorbed during training, frozen in its parameters. The source of factuality errors.
- Knowledge cutoff
- The date past which the model has no parametric knowledge. Anything that changed since is described confidently in its previous state.
- Contextual knowledge
- Facts read from the current context rather than recalled. Always prefer it for anything specific to your system.
- Attention budget
- Each token has a fixed amount of influence to distribute across the rest of the context. The weights sum to one, so context is zero-sum.
- Attention degradation
- Signal on the relationships that matter weakens as the window fills with relationships that do not. Gradual, silent, and with no error message.
- Smart zone
- The early part of a session where the agent is sharp, recalls instructions and stays in scope. Budget the zone, not the window.
- Dumb zone
- What follows: forgotten constraints, repeated mistakes, drift toward generic idiom. Same model, same prompt — only context volume differs.
- Context poisoning
- Material in the window that actively steers wrong rather than merely diluting: an abandoned approach, a corrected mistake, a stale file, an adjacent fact, an injected instruction.
- Drift
- Output pulled toward the training distribution — generic idiom rather than your conventions — as your codebase's share of the window falls. Self-reinforcing once drifted code lands.
- Sycophancy
- Agreeable output produced because agreement was rewarded in training. Ask for the case against, not for confirmation.
- Speculative generality
- Extension points, configuration and abstractions for cases that do not exist. The characteristic failure of delegated design.
- Prompt injection
- Instructions inside content the agent reads — an issue, a page, a CSV cell — being followed. Possible because there is no privileged channel.
Handoffs
- Clearing
- Ending a session and starting with an empty context. The main hygiene practice, not an admission of failure.
- Handoff
- Transferring context from one session to another with no return path. Everything that matters must be in the artefact.
- Handoff artefact
- The document that carries context across the gap: a spec, a ticket, a decisions file or a handoff note.
- Spec
- A handoff artefact describing multi-session work: goal, exclusions, constraints, decisions with reasons, acceptance criteria. Free of session-specific detail.
- Ticket
- A handoff artefact scoping exactly one session, executable by an agent with no memory of any previous one.
- Compaction
- Summarising earlier history to free room. Converts primary sources into secondary ones and drops reasons and negative results, silently.
- Autocompact
- Compaction triggered automatically as the window fills. An alarm that you are long past the smart zone, not a feature.
- Primary source
- The thing itself: the actual file, the real schema, the raw output. Complete and current, expensive to load.
- Secondary source
- An account of a primary: a README, a summary, a design doc, anything the agent summarised earlier. Cheap, lossy, and believed.
- Decisions file
- An append-only record of decisions with their reasons and dates. The reason decays faster than the decision, which is why they are kept together.
Memory and steering
- AGENTS.md
- The project's standing brief, loaded at session start. Names differ by harness; the role does not. Belongs in version control.
- Progressive disclosure
- Loading only what is needed now, with pointers to the rest. Applies to documentation, data and agent output alike.
- Context pointer
- A mention telling the agent where to look and under what condition. A pointer without a trigger is read always or never.
- Skill
- A named procedure kept out of context until a pointer activates it. For work that applies to some sessions, not most.
- Exemplar
- An existing file or diff used as the specification for new work. Communicates the conventions nobody has written down.
- Subagent
- An agent spawned by another, running in its own window. Buys a clean context; costs everything the parent knows.
- Orchestrator
- An agent holding a plan and dispatching work to others. Usually better as a shell loop with a plan file.
- The check ladder
- Four rungs of enforcement: said in the session, written in the brief, enforced by a check, impossible by construction. The skill is moving rules up.
Verification
- Automated check
- Deterministic verification: types, lint, tests. Costs no context, never tires, runs inside the agent's own correction loop.
- Check command
- The single fast command the agent runs to find out whether it broke something. Speed determines frequency, and frequency is the whole mechanism.
- Automated review
- A fresh session judging a diff it did not write. Non-deterministic, excellent at mechanical failure modes. A filter, not a gate.
- Action rate
- The fraction of an advisory channel's findings that anyone acts on. Below 40% it will be ignored within a month.
- Adversarial verification
- A second pass whose job is to refute each finding. Roughly halves false positives, which is what keeps people reading.
- Characterisation test
- A test capturing what code currently does, including what looks wrong. The only way to build a net around behaviour nobody can state.
- Property test
- A test asserting something holds for generated inputs. Covers the cases nobody thought of, which is exactly the gap generated code leaves.
- Mutation testing
- Deliberately changing code to see whether any test fails. Tells you what coverage cannot: which behaviours no test distinguishes.
- Differential testing
- Running old and new implementations against the same inputs and comparing. The strongest check available for a refactor or migration.
- Golden file
- A committed known-good output, diffed on every run. Catches serialisation-level changes that schema checks miss.
- Contract snapshot
- A generated, committed record of a published surface — API schema, exported types, event payloads. Turns an invisible break into a diff someone must approve.
- Shadow mode
- Running a new implementation alongside the old, reporting divergences, with the old one authoritative. What makes strangling a legacy system safe.
- Defect escape rate
- Defects reaching production per week or per feature. The one metric that cannot be improved by working faster.
- Review depth
- The proportion of changed lines somebody genuinely read. Falls silently under volume; nobody reports it.
- Trust calibration
- Matching review depth to check coverage and blast radius, per area. Drifts toward over-trust because failures are rare and delayed.
- Audit sample
- Properly reviewing something you approved with a skim. The only instrument that measures your own judgement rather than the system's.
Patterns of work
- Human-in-the-loop
- Pairing: watching tool calls, correcting plans, reviewing as it goes. Buys correction, costs attention.
- AFK
- Initiating and walking away. Buys throughput, requires a complete spec, stop conditions and a contained environment.
- Vibe coding
- Accepting output without reading it. Fine for throwaway work; expensive the moment the code has a second reader.
- Grilling
- Having the agent interview you, one question at a time, until the requirement is pinned down. The highest-leverage specification technique.
- Prototyping
- Building a rough version to answer one named question, then deleting it. Now cheap enough to beat another round of discussion.
- Stop condition
- A stated situation in which halting and reporting is the correct outcome. Without one, a blocked unattended agent works around the obstacle.
- The four evasions
- Widening a type, deleting an assertion, skipping a test, adding a suppression — the moves an agent makes when the honest path to a green check is blocked.
- Poison probe
- Asking the agent to list every rejected approach and every instruction still in force. Diagnoses a compromised window in one turn.
- Canary constraint
- A distinctive, checkable instruction used to measure where your own dumb zone starts.
- Eval suite
- A small fixed set of tasks from your own history, run against a model or a setup change. An afternoon once, an answer in an hour thereafter.
People and systems
- DX
- Developer experience: how well a codebase lets humans work.
- AX
- Agent experience: how well an environment lets agents perform. Mostly overlaps with DX, with searchability, locality and explicitness weighted higher.
- Hyrum's law
- With enough consumers, every observable behaviour of your system is depended on by someone, regardless of what you documented.
- Strangler pattern
- Replacing a system incrementally behind a facade, routing one piece at a time, deleting the old piece once the new one has proven itself.
- Seam
- A place where behaviour can be substituted without editing the surrounding code. Finding one is most of the work in testing legacy code.
- ADR
- Architecture decision record: a dated record of what was decided and why. The one prose format that does not go stale.
- Bus-factor question
- "What would we be stuck on if one person were unavailable for two weeks?" The answers are your runbook backlog, in priority order.