Verification is the bottleneck
Generation got a hundred times cheaper. Checking did not. Everything about how you work now follows from that asymmetry.
Two costs used to be roughly balanced. Writing a hundred lines took an hour; reading a hundred lines took ten minutes. Review was a fifth of the work and nobody optimised it.
Writing a hundred lines now takes ninety seconds. Reading them still takes ten minutes. Review has gone from a fifth of the work to nearly all of it, and almost nothing about how teams work has adjusted.
The three consequences
1. Your throughput is bounded by review, not generation. Every technique that produces more code faster is optimising a stage that is no longer the constraint. This is why module 05 keeps arriving at the same conclusion.
2. Better models make review harder, not easier. Weak output fails visibly — it does not compile, it is obviously wrong. Strong output fails subtly: it compiles, passes the tests you have, reads well, and is wrong in a way that needs attention to see. Per line, better code is more expensive to check.
3. The only real lever is moving work out of human review. Not reviewing faster — reviewing less, because a machine decided it first.
The verification stack
| Layer | Decides | Cost per run | Reliability |
|---|---|---|---|
| Types | Shape correctness | Seconds | Total, within its scope |
| Lint and custom rules | Convention, banned patterns | Seconds | Total, within its scope |
| Tests | Behaviour you thought of | Seconds to minutes | Total, within coverage |
| Property tests | Behaviour you did not think of | Minutes | High, probabilistic |
| Automated review | Mechanical smells, scope, missing requirements | A minute, cents | Partial, non-deterministic |
| Human review | Is this right, does it fit, what will it cost us | Minutes, scarce | Variable, and falls with volume |
Read the last column downward. Reliability falls as you descend, and so does throughput. The design goal is to push every decision as far up the stack as it will go, so that the bottom layer sees only what genuinely requires judgement.
The question that reorganises your week
For every defect you catch by reading a diff, ask: could a layer above have caught this?
- A swallowed error — a lint rule can catch empty and log-only catches.
- A widened type — a rule, or an
anybudget checked against a baseline. - An unhandled case — an exhaustive switch with a
nevercheck. - A changed public contract — a committed API surface snapshot.
- Scope creep — a diff-path check against the ticket.
- A test that asserts the implementation — mutation testing, or a reviewer with that single brief.
Every one of those is a thirty-minute investment that removes a category from your attention permanently. Over a few months this is the difference between review being a bottleneck and review being a spot check.
What must stay human
Push hard, but not everything moves. These stay:
- Is this the right thing to build? No layer above has your intent.
- Does it fit the system? Coherence is a judgement about a whole you hold and the checks do not.
- What will this cost in a year? Maintenance burden is invisible to every automated layer.
- Security and authorisation boundaries. Correct-looking and wrong are indistinguishable to everything above.
The point of the stack is not to remove you. It is to make sure that when you read, you are reading the four things only you can decide — rather than hunting for a missing await.
Exercise
Take every defect you found by reading a diff in the last month — from review comments, from memory, from your log. For each, write down which layer of the stack could have caught it.
Count them. Then build the single check that would have caught the most, and delete the corresponding item from your mental review checklist.
Worked solution
Twenty-three defects found by hand over four weeks, classified.
types 2 both were 'any' escaping a boundary lint / custom 9 <- the big one tests 4 real behaviour bugs, genuinely needed a test property tests 1 an off-by-one at a boundary nobody thought to test automated review 3 scope creep and one swallowed error irreducibly human 4 design fit, a bad abstraction, two "wrong thing"
4x catch block that logged and continued 2x direct process.env read outside src/config 2x new Date() used instead of the injectable clock 1x a fetch() call not going through the http client
Four of the nine were the same defect. One custom rule — ban catch blocks whose body does not rethrow, return an error value, or call a named error handler — took about forty minutes including tests and false-positive tuning.
// eslint-rules/no-silent-catch.js
// A catch must do one of: rethrow, return a Result/error value, or call
// one of our error handlers. Logging and continuing is not error handling.
message: "This catch swallows the error. Rethrow, return an error value, "
+ "or call reportError(). If continuing is genuinely correct, add "
+ "// eslint-disable-next-line no-silent-catch with a reason."The following month: zero swallowed errors reached review, and the rule fired eleven times inside agent sessions — where the agent fixed them itself, because the message said what to do. Eleven defects that never became diffs I had to read.
The four irreducibly human ones are the argument for the whole exercise. They were the most valuable findings of the month — one was an abstraction that would have cost us for a year — and I found them at the end of sessions where I had spent most of my attention on catch blocks.
Takeaways
- Review is now nearly all of the work; generation is not the constraint.
- Better models produce output that is harder per line to review, not easier.
- For every hand-caught defect, ask which layer above could have caught it, and build that.
- Reserve human attention for intent, fit, long-term cost and security — nothing above can decide those.
Check yourself
Why does a capability improvement in the model increase your review burden per line?
Obvious nonsense is rejected in seconds. Code that compiles, passes tests, reads well and is wrong in one subtle place costs a careful read. As quality rises, the proportion of defects that survive the cheap layers rises with it.
A course by Pieter Zandbergen