The Next Steps — Agent Engineering in Depth
Lessons marked free are open to read in full. The rest unlock when you buy.
01 · Mechanics You Can Reason About
What is actually happening inside a request, at the level of detail that changes decisions: attention, position, tokenisation, sampling, reasoning budgets and caching.
What a forward pass actually does
Enough mechanism to stop guessing. Not a maths lesson — a working model of the thing you are delegating to.
6 minfree 1.2
Attention is the budget
The free course told you the smart zone exists. This is why it exists, and how to predict where yours is.
5 minpaid 1.3
Position, recency and the middle
Where a fact sits in your context changes how strongly it lands. That is a lever, and most people never touch it.
5 minpaid 1.4
Tokenisation, and why code is expensive
Your context budget is denominated in tokens, and code converts to tokens far worse than prose does.
5 minpaid 1.5
Sampling, temperature and honest determinism
Where variance really comes from, which knobs help, and why reliability has to live outside the model.
5 minpaid 1.6
Reasoning budgets: what effort actually buys
Extended thinking is real and measurable. It is also the most commonly misapplied setting available to you.
5 minpaid 1.7
The prefix cache and the shape of your context
Providers cache the unchanged start of your context. Designing around that changes both your bill and your habits.
4 minpaid 1.8
Tool calling under the hood
A tool call is generated text that your harness chose to execute. Everything odd about tool use follows from that.
5 minpaid 1.9
Drift to the training distribution
Why agents write code that looks like a tutorial rather than like your repository, and what actually stops it.
6 minpaid 1.10
What a better model actually changes
Which of your problems a model upgrade solves, which it makes worse, and how to decide whether to switch.
4 minpaid
02 · Context Engineering
Treating the context window as something you design rather than something that fills: retrieval, session shapes, poisoning, memory systems, compaction and very large inputs.
The context budget as a design artefact
Decide what goes in the window before the session starts, the same way you would decide an interface before writing the class.
6 minfree 2.2
Retrieval that works
Why grep and the type system usually beat embeddings for code, and how to build a retrieval habit that scales.
5 minpaid 2.3
Designing the opening message
The highest-leverage two hundred tokens you will write. A structure, and the failure each part prevents.
5 minpaid 2.4
Five session shapes
Explore, decide, implement, verify, repair. Each needs different context, a different budget and a different ending.
5 minpaid 2.5
Context poisoning
Some things in a window do not just waste space — they actively steer the agent wrong. Learn to spot them.
6 minpaid 2.6
Memory systems, honestly
What persistent memory actually is, the two ways it fails, and how to get the benefit without the failure.
4 minpaid 2.7
Compaction, deliberately
Automatic compaction is a lossy accident. The same operation performed on purpose is one of the better tools you have.
5 minpaid 2.8
Very large inputs
A 40,000-line file, a 200 MB log, a database with 300 tables. Strategies that do not involve loading them.
6 minpaid 2.9
Monorepos and multi-repo context
When the thing you are changing spans several packages or several repositories, context assembly becomes the hard part.
5 minpaid 2.10
Instrumenting context
Three measurements that turn context management from a feel into a number, and what to do with each.
5 minpaid
03 · Building Your Own Tooling
Stop working around your harness and start extending it: tools designed for agents, check commands as products, commands, skills, hooks, small MCP servers and scripted runs.
Why you should build tools
The gap between a good agent setup and an average one is mostly tooling, and it is cheaper to build than it has ever been.
5 minfree 3.2
Designing a tool for an agent
Agent-facing tools have different design rules from human-facing ones. Six of them, with the failure each prevents.
5 minpaid 3.3
The check command as a product
The single highest-leverage thing in your repository. Treat it as something you design, not something that accumulated.
5 minpaid 3.4
Custom commands and prompt templates
Your rituals, saved. The cheapest tooling you can build, and the one most people skip.
5 minpaid 3.5
Writing a skill that survives
Procedures rot faster than conventions. How to write one that fails loudly when it goes stale.
6 minpaid 3.6
Hooks and lifecycle automation
Code that runs at fixed points in the agent loop. Deterministic enforcement that does not depend on attention at all.
3 minpaid 3.7
Writing a small MCP server
When your agent needs to reach something outside the filesystem, and how to do it without taxing every session.
6 minpaid 3.8
Scripting the agent
Non-interactive invocation turns an agent into a component you can put in a pipeline, a hook, or CI.
5 minpaid 3.9
A repo-local agent toolkit
Everything in this module, assembled into one directory that travels with the project.
5 minpaid 3.10
Maintaining your tooling
Tools rot. The maintenance habits that keep a toolkit worth its permanent cost.
5 minpaid
04 · Specification and Design
The discipline that decides whether unattended work is possible at all: interviewing yourself, acceptance criteria that bite, interface-first design, and specifying data, migrations and ambiguity.
The specification is the product
Once generation is cheap, the scarce artefact is a precise statement of what to generate. That is now most of the job.
6 minfree 4.2
Grilling at depth
The interview technique, done properly: how to steer it, what to do when you do not know, and where it fails.
6 minpaid 4.3
Design concepts versus deliverables
The shared understanding of what you are building is a separate thing from any document, and it is what actually has to be right.
6 minpaid 4.4
Acceptance criteria that bite
Most criteria are restatements of the requirement. Useful ones are observations that could fail.
6 minpaid 4.5
Interface-first specification
Write the signatures before the prose. They are the most compressed, least ambiguous specification available.
6 minpaid 4.6
Specifying data and migrations
The part of a spec that is hardest to reverse, and the one most often left to the implementation session.
6 minpaid 4.7
Specifying for unattended execution
The extra things a spec needs when nobody is watching, and how to find out whether yours has them.
6 minpaid 4.8
Ambiguity you cannot resolve
Sometimes the answer genuinely is not known yet. Specify that honestly rather than pretending it away.
6 minpaid 4.9
Keeping specs alive
A spec that stops matching reality becomes a confident secondary source. How to keep them true or kill them.
6 minpaid
05 · Orchestration and Multi-Agent Work
Running more than one agent without producing more than one codebase: subagents, parallel worktrees, orchestrator patterns, shared state, failure handling and knowing when to stop.
When more agents help, and when they do not
Parallelism is real but narrow. A decision procedure, before you build anything elaborate.
6 minfree 5.2
Subagents in depth
The highest-value multi-agent pattern, and the four instruction failures that spoil it.
5 minpaid 5.3
Parallel work with worktrees
The practical mechanics of running several agents at once without them destroying each other.
4 minpaid 5.4
The orchestrator pattern
One agent holding the plan and dispatching work. When it earns its keep, and the two ways it fails.
6 minpaid 5.5
Agent-to-agent review
The most valuable multi-agent pattern there is, and the specific ways it degrades into noise.
5 minpaid 5.6
Pipelines, fan-out and fan-in
Three composition shapes, what each is good for, and the reduce step that most people get wrong.
5 minpaid 5.7
Shared state between agents
When several agents need to agree on something, the answer is a file with rules — not a conversation.
5 minpaid 5.8
Failure handling in multi-agent runs
What to do when one of twelve tasks fails, and why the obvious answers are wrong.
6 minpaid 5.9
The economics of orchestration
What parallel agents actually cost in money, latency and attention — and where the real ceiling is.
5 minpaid 5.10
When to stop orchestrating
The signals that your setup has become more work than the work, and what to do about it.
5 minpaid
06 · Verification and Evaluation
The bottleneck, taken seriously: check ladders, property-based testing, a personal eval suite, reviewing at volume, detecting subtle breakage, and calibrating how much to trust.
Verification is the bottleneck
Generation got a hundred times cheaper. Checking did not. Everything about how you work now follows from that asymmetry.
3 minfree 6.2
Climbing the check ladder
Four rungs from "I said so" to "it cannot happen". A method for moving any rule up.
5 minpaid 6.3
Property-based testing
Tests for the cases you did not think of. Especially valuable when the implementation was generated rather than reasoned about.
3 minpaid 6.4
Building an eval suite for your repository
An afternoon once, an answer in an hour on every future model release. The best-value thing in this module.
4 minpaid 6.5
Automated review that earns its place
How to keep the hit rate high enough that people still read it after a month.
4 minpaid 6.6
Reviewing at volume
Concrete techniques for reading more diffs without reading them worse.
3 minpaid 6.7
Detecting subtle breakage
The defects that pass every check and read fine. Five techniques for finding them before your users do.
5 minpaid 6.8
Regression safety nets
Making the first occurrence of a defect the last. The discipline that compounds fastest.
6 minpaid 6.9
Metrics that mean something
Four numbers worth tracking, three that will mislead you, and what to do when they move.
3 minpaid 6.10
Trust calibration
How much to check, when, and how to notice that your calibration has drifted.
4 minpaid
07 · Legacy and Large Codebases
What changes when the code is old, large, undocumented or frightening: characterisation, seams, strangler patterns, undocumented behaviour and the politics of touching things nobody owns.
What legacy actually changes
Every assumption in the previous six modules quietly depends on things a legacy codebase does not have.
5 minfree 7.2
Characterisation first
Before you change untested code, capture what it currently does — including the parts that look wrong.
5 minpaid 7.3
Finding seams
A seam is a place you can change behaviour without editing the code around it. In legacy work, finding one is most of the job.
5 minpaid 7.4
The strangler pattern with agents
Replacing a system incrementally while it keeps running — and where cheap generation changes the calculus.
6 minpaid 7.5
Code archaeology
Using history, not just source, to find out why a line exists. The single best defence against deleting something load-bearing.
6 minpaid 7.6
Navigating very large codebases
A million lines, four hundred modules, and an agent that can hold about two files. Strategies that scale.
5 minpaid 7.7
Undocumented behaviour as a hazard
Hyrum's law, applied: every observable behaviour of your system is depended on by somebody.
6 minpaid 7.8
The politics of touching things nobody owns
The technical part is often the easy part. Who to ask, what to promise, and how to avoid becoming the owner.
6 minpaid
08 · Security, Safety and Supply Chain
The risks that come with delegating code and command execution: prompt injection, secrets, dependencies, permissions, data handling and reviewing security-sensitive diffs.
The threat model
What is actually different about running an agent in your development environment, stated precisely.
5 minfree 8.2
Prompt injection in practice
Where untrusted text actually reaches your agent, what it can do, and the defences that do not depend on the model.
4 minpaid 8.3
Secrets and credentials
Keeping secrets out of context, out of diffs, and out of the agent's reach entirely.
3 minpaid 8.4
Supply chain
Agents add dependencies casually and sometimes invent them. Both are worth a gate.
5 minpaid 8.5
Permissions and sandboxing in practice
Concrete configurations for three levels of autonomy, and the approval-fatigue trap in between.
4 minpaid 8.6
Reviewing security-sensitive code
Why generated authorisation and validation code needs a different review, and what to look for.
6 minpaid 8.7
Data handling and privacy
What leaves your machine when you use an agent, and the rules that keep it defensible.
6 minpaid 8.8
A posture that fits
Assembling everything in this module into three concrete postures, and choosing one honestly.
5 minpaid
09 · Cost, Performance and Operations
Running this sustainably: where the money goes, latency and flow, model selection, budgets, failure modes in production and the operational habits that keep it working.
Where the money actually goes
A cost model precise enough to act on, and the three line items that dominate every bill.
5 minfree 9.2
Latency and flow
The cost nobody measures: waiting, context-switching, and what an interrupted engineer actually costs.
5 minpaid 9.3
Model selection in practice
Matching model to task, when switching is worth it, and how to decide with evidence rather than vibes.
5 minpaid 9.4
Budgets and guardrails
Spending controls that catch runaway usage without making ordinary work annoying.
3 minpaid 9.5
Agents in production workflows
Putting an agent in a path that runs without you: what to allow, what to forbid, and how to fail.
6 minpaid 9.6
Operational habits
The small recurring practices that keep everything in this course working six months from now.
5 minpaid
10 · Teams, Process and Craft
Making this work with other people: adoption, shared conventions, code review at volume, hiring and mentoring, and what happens to your own skill.
Adoption without a mandate
How this actually spreads through a team, and the three ways organisations get it wrong.
6 minfree 10.2
Shared conventions at team scale
One brief, several people, several agents. Keeping conventions consistent when nobody owns them.
5 minpaid 10.3
Code review at team scale
Four people generating at volume saturates review in a fortnight. What to change before that happens.
5 minpaid 10.4
Mentoring and junior engineers
The hardest question in this course: how people learn to engineer when the doing part is delegated.
6 minpaid 10.5
Hiring and evaluating engineers
What to assess when writing code is no longer the scarce skill, and how to run an interview that finds it.
5 minpaid 10.6
What happens to your own craft
Which of your skills atrophy, which become more valuable, and what to do about it deliberately.
6 minpaid
11 · Across the Stack
What actually differs by domain: frontend, data work, infrastructure, APIs, mobile and machine learning each break the general advice in their own specific way.
Why domain matters
The general process is right. The verification layer underneath it is completely different in each domain, and that changes everything.
4 minfree 11.2
Frontend work
Where the check stops at "it rendered" and the interesting questions all start after that.
5 minpaid 11.3
Data work
Where the code can be perfect and the numbers still wrong, and nothing tells you.
3 minpaid 11.4
Infrastructure and configuration
Where the check tells you what you intend to do and nothing tells you what will happen.
5 minpaid 11.5
APIs and contracts
Anything with consumers you do not control changes the rules completely.
5 minpaid 11.6
Mobile and desktop
Long feedback loops, irreversible releases, and users who never upgrade.
5 minpaid 11.7
Machine learning code
Where the pipeline can be correct and the result still worse, and "it runs" means almost nothing.
5 minpaid 11.8
Crossing domains in one change
A feature that touches frontend, API, database and infrastructure needs a different decomposition.
6 minpaid
12 · Debugging and Incident Response
The skill agents help with least and cost most when done badly: hypothesis discipline, instrumentation, bisecting, production incidents, and the failures that only happen at scale.
Why debugging is different
Every other task starts from a specification. Debugging starts from a contradiction, and that changes the whole method.
5 minfree 12.2
The four-step structure
Reproduce, hypothesise, discriminate, fix — with the agent doing the work at each step and you deciding between them.
5 minpaid 12.3
Instrumentation over speculation
The observation is almost always cheaper than the guess. Building it is the thing to delegate.
5 minpaid 12.4
Regressions and bisecting
When it used to work, the history contains the answer and the search is mechanical.
5 minpaid 12.5
Production incidents
Where an agent helps, where it is dangerous, and the discipline that stops it making a bad hour worse.
6 minpaid 12.6
Intermittent and timing bugs
The category where the general method needs most adjustment, because you cannot reliably reproduce.
5 minpaid 12.7
Performance investigation
Where intuition is least reliable and measurement is cheapest, and where agents guess most confidently.
5 minpaid 12.8
When to stop and read it yourself
The signals that further prompting will not help, and what to do instead.
6 minpaid
13 · Architecture for Agent-Assisted Work
Design decisions reconsidered: what makes a codebase easy for agents, which classic trade-offs have shifted, and how to build systems that stay workable as generation gets cheaper.
Which trade-offs actually shifted
Most architectural advice survives unchanged. Four trade-offs genuinely moved, and knowing which matters more than any new principle.
5 minfree 13.2
Module boundaries that hold
Boundaries that are enforced rather than intended, and why that distinction now matters more.
5 minpaid 13.3
Designing for verification
Structure the system so that correctness is checkable, because checkability is what determines how much you can delegate.
5 minpaid 13.4
Documentation that survives
Most documentation is a stale secondary source. The kinds that stay true, and how to make the rest fail loudly.
5 minpaid 13.5
When to let an agent design
Which design decisions to delegate, which to keep, and how to use an agent as a critic rather than an author.
6 minpaid 13.6
Evolution and deletion
Cheap generation makes codebases grow faster. Deletion has become the scarce discipline.
6 minpaid
14 · Research and Unfamiliar Territory
Using an agent to learn: evaluating libraries, reading unfamiliar code and papers, picking up a new language, and the specific ways confident wrongness bites hardest when you cannot check.
Learning without ground truth
The situation where confident wrongness costs most: you cannot evaluate the answer, because not knowing is why you asked.
5 minfree 14.2
Evaluating a library or tool
A repeatable process for deciding whether to adopt something, using an agent for the parts it is good at.
4 minpaid 14.3
Reading unfamiliar code
Someone else's repository, a dependency's internals, a codebase you have just joined.
6 minpaid 14.4
Learning a new language or framework
Where agents accelerate most and mislead most, and how to get the first without the second.
6 minpaid 14.5
Reading papers and specifications
Dense primary sources are exactly where an agent helps, provided the source stays in the window.
5 minpaid 14.6
Synthesis across many sources
When the question needs ten documents, the fan-in problem returns — and the reduce step is where it goes wrong.
5 minpaid
15 · Worked Projects
Six complete pieces of work, start to finish, with every decision and every mistake shown: a feature, a migration, an incident, a rescue, a greenfield service and a team rollout.
How to read these
Six projects, shown with the wrong turns left in. What to look for, and what not to copy.
5 minfree 15.2
Project: a feature, end to end
Scheduled CSV exports, from a one-line request to production, with every session shown.
6 minpaid 15.3
Project: a large migration
Ninety-one files from one validation library to another, exemplar-driven, with the batch failures shown.
3 minpaid 15.4
Project: a production incident
Forty minutes of checkout failures, with the moment the discipline slipped and what it cost.
6 minpaid 15.5
Project: rescuing a bad session
Two hours into a session that has gone wrong. Recognising it, and getting out with what is worth keeping.
6 minpaid 15.6
Project: a greenfield service
Day one of a new service, building the verification infrastructure before the first feature.
6 minpaid
16 · Beyond Code
The rest of an engineer's job: writing, planning, estimating, technical decisions, runbooks, communicating with non-engineers, and knowing when not to use an agent at all.
The rest of the job
Writing code was never most of engineering. Now that it is cheap, the other parts are most of it explicitly.
5 minfree 16.2
Technical writing
Documents other people have to act on: proposals, postmortems, updates, and the critique loop that improves them.
6 minpaid 16.3
Planning and estimation
Where agents are genuinely poor, why, and the narrow ways they still help.
6 minpaid 16.4
Runbooks and operational documentation
The documentation nobody writes, which is exactly the kind an agent should write.
6 minpaid 16.5
Making technical decisions
Using an agent to decide better without letting it decide, and recording the result so it survives.
6 minpaid 16.6
Communicating with non-engineers
Translating technical reality for people who will act on it, without making it false.
6 minpaid 16.7
When not to use one at all
The situations where reaching for an agent makes the work worse, and recognising them early.
6 minpaid
Unlock all 128 lessons
One payment, lifetime access, every future revision included.
Buy for €149,00