Why domain matters
The general process is right. The verification layer underneath it is completely different in each domain, and that changes everything.
Everything so far has assumed a particular shape of work: source files, a type checker, a test suite, a diff. That shape holds for backend application code and holds progressively less well as you move away from it.
The process does not change. Small tasks, curated context, a spec, a check, a review. What changes is what the check can be — and since the check is what makes delegation safe, a domain with a weak check is a domain where you delegate less until you have built one.
The verification landscape
| Domain | What a check can decide | What it cannot |
|---|---|---|
| Backend application code | Almost everything that matters | Design fit, long-term cost |
| Frontend | Types, rendering, accessibility tree, some interaction | Whether it looks right, whether it feels right |
| Data pipelines | Schema, row counts, invariants on a sample | Whether the numbers are correct |
| Infrastructure | Plan diffs, policy rules, syntax | What happens when it is applied |
| APIs with consumers | Contract snapshots, schema compatibility | Whether a consumer relies on something undocumented |
| Mobile | Types, unit tests, some UI tests | Device behaviour, store review, upgrade paths |
| Machine learning | Shapes, pipeline correctness | Whether the model is better |
Read the third column. In backend work it contains two things that are genuinely rare. In infrastructure it contains the entire point of the change.
The rule this produces
Delegation depth should follow verification strength, per domain, not per person. The same engineer, in the same day, should work almost hands-off on a well-tested backend module and read every line of a Terraform plan — not because one is harder but because one has a check that decides and the other does not.
The three moves that work everywhere
Before the domain-specific lessons, three things transfer:
- Find the strongest check the domain allows, and make it fast. In frontend that might be an accessibility-tree snapshot; in data, a set of invariants; in infrastructure, a plan diff with policy rules.
- Make the gap explicit. Write down what your check cannot decide, so you know exactly what your own review is for. Half of bad delegation is not knowing where the check stops.
- Shorten the feedback loop before anything else. Every domain has a slow loop that can be made faster, and speed determines whether the agent self-corrects or hands you unverified work.
What follows
The rest of this module is one lesson per domain. Each one names the specific way the general advice breaks there, the strongest available check, and the practices that work. Read the ones you work in; skim the rest, because the failure modes rhyme and knowing the shape helps when you cross a boundary.
Exercise
For the domain you work in most, fill in your own row of the table honestly: what can your checks actually decide, and what is left entirely to you?
Then write the second column as a list and put it in your project brief, because that list is your review checklist — and most teams have never written it down.
Worked solution
The row filled in for a React frontend, honestly.
TypeScript props, state shapes, event handler signatures
ESLint + a11y plugin missing labels, bad ARIA, non-interactive
elements with handlers
Vitest + testing-lib does the right text appear, do handlers fire,
does the loading state render
Playwright (12 flows) the critical paths still work end to end
Chromatic a visual diff, which a human still judges- whether the layout is right at any width we do not screenshot - whether the interaction feels responsive - whether the empty state is helpful or bleak - whether the error message tells the user what to do - whether it works with a screen reader in practice (the a11y lint catches structure, not usability) - whether it degrades acceptably on a slow connection - whether this is consistent with the rest of the product - whether animation is doing something or just happening
Eight items, none of them exotic, and writing them down changed how review worked immediately. Before, reviewing a frontend PR meant reading the diff and looking at a screenshot. Now it means reading the diff for the things the checks miss, and opening the preview deployment with those eight questions in mind.
Two of the eight became checks within a month, which is the other benefit of writing the list: it is a backlog. A Playwright test that throttles the network caught a loading state that had never been visible in development, and an axe run in CI caught a real focus-order problem. Six remain human, and now everyone knows which six.
Takeaways
- The process is domain-independent; the strength of the check is not.
- Delegation depth should follow verification strength per domain, not per person.
- Write down what your checks cannot decide — that list is your actual review checklist.
- The list is also a backlog: some items become checks once you can name them.
Check yourself
Why should the same engineer delegate more on backend code than on infrastructure changes?
Delegation is safe in proportion to what a check can decide without you. A green backend test suite settles most questions; a valid Terraform plan settles syntax and shows an intended diff, but not the consequences of applying it. The gap is what your review must cover.
A course by Pieter Zandbergen