How to read these
Six projects, shown with the wrong turns left in. What to look for, and what not to copy.
The rest of this module is six complete pieces of work. They are written with the mistakes left in, because a worked example that goes smoothly teaches nothing you could not get from the technique lessons.
What each one shows
| Project | The thing it demonstrates |
|---|---|
| A feature, end to end | The full chain: grill, spec, decompose, build, verify, retrospect |
| A large migration | Exemplar-driven batch work, and where it breaks |
| A production incident | Discipline under time pressure, and what it costs to lose it |
| Rescuing a bad session | Recognising and recovering from a poisoned window |
| A greenfield service | Building the verification infrastructure before the features |
| A team rollout | Adoption in practice, over eight weeks |
What to look for
The decision points, not the prompts. The prompts are illustrative and will date. The decisions — when to clear, when to stop delegating, when to read it yourself — are the transferable part.
Where the process was skipped and what it cost. Each project contains at least one place where a step was skipped under pressure. Those are the most instructive passages.
The numbers. Token counts, session counts, review times, defects found. They calibrate expectations better than any description, and they let you compare against your own.
What not to copy
- The exact commands. Your harness, language and tooling differ.
- The specific structure. Five sessions for a feature is not a rule; it is what that feature needed.
- The ratios. "One hour of spec for four hours of build" is a description of one case, not a target.
A caution about worked examples
Every case study reads as more orderly than it was. The sequence is clear in hindsight, the dead ends are compressed, and the reasoning is reconstructed. Real work is messier than any account of it, including these.
So treat them as demonstrations of shape rather than as scripts. The useful question while reading is not "what did they do" but "at this point, what would I have done, and why?" — and then to notice where your answer differs and whether you can say why.
The one thing common to all six
Each project spends more effort than feels necessary on deciding what to build and checking what came back, and less than feels necessary on producing it. That inversion is the whole course, and it is more visible in a worked example than in any argument for it.
Exercise
Before reading the next five lessons, write down how you would approach each of the six situations in the table above, in two or three lines each.
Then compare as you read. The places where you differ are worth thinking about in both directions — some of your instincts will be better suited to your situation than anything here.
Worked solution
The comparison exercise, done by an engineer who did it before reading the module.
"A feature, end to end"
predicted: spec, then build, then review.
differed: had not separated grilling from spec-writing, or explore
from decide. Four sessions where they expected two.
Their comment afterwards: "the grilling session is the one
I would have skipped and it is the one that changed the
design."
"A large migration"
predicted: exemplar + batch, roughly right.
differed: had not planned for the batch to fail on a subset, or for
three identical failures to mean a spec gap rather than
three special cases.
"A production incident"
predicted: get the agent to help diagnose fast.
differed: had not considered that mitigation should be entirely
human. Their comment: "I would absolutely have let it
propose a rollback command and run it. Under pressure I
would not have read it carefully."
"Rescuing a bad session"
predicted: start over.
differed: right instinct, but had not thought about what to carry -
specifically that ruled-out approaches need their reasons
carried with them or they come straight back.
"A greenfield service"
predicted: build the feature, add tests as they go.
differed: the module builds the check command, the brief and the
first lint rules before the first feature. Their comment:
"this feels backwards and I can see why it is not."
"A team rollout"
predicted: show people, write a guide.
differed: the module puts the toolkit in the repository and says
nothing. They had not considered that the repository
teaches faster than a document.Five of six predictions were reasonable and incomplete in the same direction: each skipped a preparatory step that felt like overhead. That is the most common pattern in this exercise, and it is worth knowing about yourself before reading, because it is much easier to notice in your own predictions than in someone else's account.
Takeaways
- Read for the decision points, not the prompts — the prompts will date and the decisions will not.
- The passages where the process was skipped are the most instructive ones.
- Every case study reads as more orderly than the work was, including these.
- Ask "what would I have done here?" before reading on, and notice where you differ.
Check yourself
What is the most common way people's predictions differ from these worked examples?
The consistent pattern is compressing the front of the process: going straight to building rather than separating out interviewing, deciding, and building the checks. Those steps feel like preamble and are where most of the leverage is, which is exactly why they get skipped.
A course by Pieter Zandbergen