Reading the diff
Review is the bottleneck now. Here is how to read agent-written code fast without becoming a rubber stamp.
Generation got cheap. Review did not. If you read agent diffs the way you read a colleague's pull request — linearly, top to bottom — you will be the slow part of your own workflow, and you will still miss things, because agent diffs fail differently from human ones.
How agent code fails differently
- Plausible-looking invention. A helper that does not exist, called correctly. Humans rarely do this; models do it constantly.
- Scope creep. Three files you asked for, plus a "while I was here" refactor in a fourth.
- Convention drift. Correct code written in a style that is not yours, which compounds across sessions.
- Tests that assert the implementation. A test written after the code, asserting exactly what the code does, including the bug.
- Silent error swallowing. A
try/catchthat logs and continues, added to make something pass.
A four-pass review
- Shape (10 seconds).
git diff --stat. Are the files the ones you scoped? Is the line count roughly what you expected? A 400-line diff for a 40-line fix is a stop signal before you read a single line. - Danger (30 seconds). Grep the diff for your own danger list:
catch,any,eslint-disable,TODO,sleep,skip, credentials, new imports. These are where agent diffs hide their compromises. - Tests first (2 minutes). Read the test changes before the implementation. Ask: would this test fail if the implementation were wrong? A test that mirrors the implementation line for line is worse than no test, because it blocks future refactors while proving nothing.
- Implementation (the rest). Now read the code, with the question "does this do what the test claims" rather than "is this how I would write it".
git diff --stat git diff | grep -nE '\b(catch|any|TODO|eslint-disable|@ts-ignore|sleep)\b' git diff -- '*test*' '*spec*' git diff -- . ':(exclude)*test*' ':(exclude)*spec*'
Make the agent help you review
The agent that wrote the code is a poor reviewer of it — it is in the same context and shares the same assumptions. A fresh session with only the diff is genuinely useful:
Here is a diff. You did not write it and know nothing about the intent beyond this description: [one line]. List anything that: changes behaviour outside the described scope, swallows an error, weakens a type, or adds a test that would pass even if the logic were wrong. Do not comment on style.
This is automated review. It is not a substitute for your eyes — it makes non-deterministic judgements — but it is very good at catching the mechanical failure modes above, which is exactly the part of review that is boring and therefore skipped.
Try it
Build your own danger-grep as a shell alias, tuned to your language and the mistakes you have actually seen. Run it on the last three agent commits in your repo. Most people find at least one thing.
Takeaways
- Check diff shape before reading any code — scope creep is visible in the stat line.
- Read the tests first and ask whether they would fail if the code were wrong.
- A fresh agent reviewing only the diff catches mechanical problems the author agent cannot see.
Why ask a different session to review, rather than asking the same agent to check its work?
Because the authoring session shares every assumption that produced the bug, and its context is full of reasons the code is right. A fresh session with only the diff has none of that priming — different context, genuinely different judgement.
A course by Pieter Zandbergen