Review at volume
When generation is cheap, review is the whole job. Layer it so your attention lands where it matters.
The bottleneck has moved. You can now produce more code per day than you can carefully read, which means the only real question is how to allocate the reading you can do.
Three layers, cheapest first
| Layer | Catches | Cost |
|---|---|---|
| Automated check types, lint, tests, custom rules | Deterministic violations. Never tires, never misses one it can express. | Seconds. Build once. |
| Automated review a fresh agent reading the diff | Scope creep, swallowed errors, weakened types, tests that assert the implementation. | Cents and a minute. |
| Human review you | Is this the right thing? Does it fit the system? What will this cost us in a year? | Your scarcest resource. |
The discipline is to never spend layer three on something layer one or two could have caught. Every time you find a lint-able problem by reading, that is a signal to write the rule.
A usable automated review prompt
Review this diff against the ticket below. You did not write it. Report only: (1) behaviour changes outside the ticket's stated scope, (2) swallowed or ignored errors, (3) weakened types or added suppressions, (4) tests that would still pass if the logic were wrong, (5) anything the ticket required that is missing. For each finding: file, line, one sentence, and how confident you are. Do not comment on style or naming. If you find nothing, say so. --- TICKET --- [paste ticket] --- DIFF --- [paste diff]
Constraining the categories is what makes this useful. An open-ended "review this" produces fifteen style opinions and buries the one real finding.
Where your attention actually belongs
- The boundaries. Public APIs, schema changes, anything other code depends on. Internal implementation can be rewritten later; boundaries cannot.
- Anything security- or money-adjacent. No check catches "correct-looking and wrong" here.
- The tests, always. They are what lets you skim implementations in future.
- The stuff you cannot check. Whether this is the right design, whether it fits, whether it will still make sense to the next person.
What to do when the diff is too big
Do not read it worse. Send it back. A diff too large to review carefully is a process failure upstream — the ticket was too big, or the scope was not enforced — and reviewing it superficially converts that into a code failure you will pay for later. Split it into commits, review them in sequence, or regenerate from smaller tickets.
Watch out
Automated review is non-deterministic and will miss things, including obvious things. It reduces the volume you have to read carefully; it does not remove the requirement. Treating it as a gate rather than a filter is how a team discovers, six months in, that nobody has actually read the code.
Try it
Run the automated review prompt over your last three agent commits. For every real finding, ask whether a check could have caught it — and write one of those checks today.
Takeaways
- Checks, then automated review, then you — never spend the expensive layer on cheap problems.
- Constrain the review agent to specific categories or it produces style noise.
- Spend human attention on boundaries, security, tests, and design fit.
- Send back a diff that is too big instead of reading it badly.
You caught a swallowed error by reading a diff. What is the follow-up action?
Write the rule. A lint rule banning empty or log-only catch blocks turns that class of finding into a layer-one catch forever, and hands your attention back for the judgement calls only you can make. Every human catch of a mechanical problem is a missing check.
A course by Pieter Zandbergen