Agent Engineering
Reference

Glossary

Shared vocabulary for the whole course. Precision here is not pedantry — most bad agent sessions come from confusing two of these terms with each other.

The model

Model
The parameters. Stateless, and does next-token prediction and nothing else. Everything that feels like memory or agency comes from the layer above it.
Inference
Running a trained model to produce output. What happens on every request to a model provider.
Token
The atomic unit a model reads and writes. Roughly four characters of English; fewer for code, far fewer for dense JSON.
Effort
A dial controlling how much internal reasoning happens before the model answers. High for deciding, low for doing.
Non-determinism
The same input can produce different output, from sampling during generation and from how the provider serves the request.
Input tokens
What you send on each request. Cheap per token, enormous in volume, and re-sent in full every turn.
Output tokens
What the model generates. Priced higher per token, but usually far fewer of them.
Prefix cache
Provider-side storage of the unchanged start of your context, billed at a discount. Makes stable context nearly free and mid-session instruction edits expensive.

Sessions and context

Harness
Everything around the model that turns it into an agent: system prompt, tools, permission model, and context-window management. Most tool comparisons are harness comparisons.
Agent
A model harnessed with tools, a system prompt, and a context window, taking turns with a user.
Context window
Everything the model sees on one request. Finite, model-specific, and re-sent in full every time.
Context
The relevant information the agent has right now. Not the same as the window: a full window can hold terrible context.
Session
One bounded run of interaction. Starts empty, accumulates, ends when cleared. The real unit of agent work.
Turn
One user message plus everything the agent does before handing control back — which may be dozens of tool calls.
System prompt
Instructions the harness prepends to every request. The agent’s standing brief, before your project’s own.
Stateless
Carries nothing forward between requests. The model is stateless; the appearance of memory is the harness re-sending history.

Tools and environment

Tool
A function the harness exposes for the agent to call: read, write, search, run a command.
Tool result
What the harness sends back after executing a call. It is appended to the context window permanently, which is why verbose output is expensive.
MCP
A protocol for plugging external tool servers into a harness. Each connected server’s definitions sit in every session’s system prompt, so connect deliberately.
Permission mode
Which tool calls run without asking you. Pre-approve the frequent safe ones so the remaining prompts get read.
Sandbox
An isolated environment the agent runs inside — container, VM, worktree, restricted shell. Lowers blast radius so autonomy can be raised.
Environment
The world the agent acts on, perceived only through tool results. Its legibility is your design problem.

Failure modes

Hallucination
Confidently-wrong output. Two flavours: factuality (invented from training) and faithfulness (drifted from loaded context).
Parametric knowledge
What the model absorbed during training, frozen in its parameters. The source of factuality errors.
Knowledge cutoff
The date past which the model has no parametric knowledge. Libraries released after it are fabrication traps.
Contextual knowledge
Facts the agent reads from its current context rather than recalls. Always prefer it for anything specific to your code.
Attention budget
Each token has a finite amount of influence to spread over the rest of the context. More tokens means a thinner slice each.
Attention degradation
Signal on the relationships that matter weakens as the session fills with relationships that do not.
Smart zone
The early part of a session where the agent is sharp, recalls instructions, and stays in scope. Budget the zone, not the window.
Dumb zone
What follows: forgotten constraints, repeated mistakes, confident claims contradicting loaded files. No error message marks the boundary.
Sycophancy
Agreeable output produced because agreement was rewarded in training. Ask for the case against, not for confirmation.

Handoffs

Clearing
Ending a session and starting with an empty context. The main hygiene practice, not an admission of failure.
Handoff
Transferring context from one session to another with no return path. Everything that matters must be in the artifact.
Handoff artifact
The document that carries context across the gap: a spec, a ticket, or a handoff note.
Spec
A handoff artifact describing the objective of multi-session work, free of session-specific detail. Goal, exclusions, constraints, decisions, acceptance.
Ticket
A handoff artifact scoping exactly one session, executable by an agent with no memory of any previous one.
Compaction
An in-memory handoff: earlier history is summarised to seed a continued session. Converts primary sources into secondary ones, silently.
Autocompact
Compaction triggered automatically as the window fills. Treat it as an alarm that you are long past the smart zone.
Primary source
The thing itself: the actual file, the real schema, the raw failure output. Complete and current, but expensive to load.
Secondary source
An account of a primary: a README, a summary, a design doc, anything the agent summarised earlier. Cheap, lossy, and believed.

Memory and steering

AGENTS.md
The project’s standing brief, loaded at session start. Names differ by harness; the role does not. Check it into the repo.
Progressive disclosure
Loading only what is needed now, with pointers to the rest. Keeps the always-loaded context small.
Context pointer
A mention in one document telling the agent where to look, and under what condition. A pointer without a trigger is read always or never.
Skill
A teachable procedure bundled as a unit and kept out of context until a pointer activates it. For work that applies to some sessions, not most.
Subagent
An agent spawned by another via a tool call, running in its own context window. Buys a clean window; costs everything the parent knows.

Patterns of work

Human-in-the-loop
Pairing: watching tool calls, correcting plans, reviewing as it goes. Buys correction, costs attention.
AFK
Initiating and walking away. Buys throughput, requires a complete spec and a contained environment.
Vibe coding
Accepting output without reading it, treating the diff as opaque. Fine for throwaway work; expensive the moment the code has a second reader.
Grilling
Having the agent interview you, one question at a time, until the requirement is genuinely pinned down. The highest-leverage technique in the course.
Prototyping
Building a rough version to answer one named question, then deleting it. Now cheap enough to beat another round of discussion.
Automated check
Deterministic verification: types, lint, tests. Costs no context, never tires, runs inside the agent’s own correction loop.
Automated review
A fresh agent judging another agent’s diff. Non-deterministic, but excellent at the mechanical failure modes. A filter, not a gate.
Human review
You reading the diff. The scarcest layer — spend it on boundaries, security, tests, and design fit.
DX
Developer experience: how well a codebase lets humans work.
AX
Agent experience: how well an environment lets agents perform. Mostly overlaps with DX, with searchability, locality, and explicitness weighted higher.

More of this vocabulary

The Next Steps uses a fuller version of this glossary — 97 terms — and it is free to read there without buying anything.

Open the full glossary →

A course by Pieter Zandbergen