When more agents help, and when they do not
Parallelism is real but narrow. A decision procedure, before you build anything elaborate.
Multi-agent setups are the most over-sold idea in this field and the most under-used where they actually work. The difference is entirely about what the second agent has that the first one does not.
The only two things a second agent buys
1. A separate context window. Work that would have polluted your window happens elsewhere, and a summary comes back. That is the whole of the subagent argument, and it is a good one.
2. Independence from your reasoning. A reviewer that never saw the implementation conversation does not share its blind spot. This is why review must be a separate session, and it is the strongest argument in the whole module.
Everything else attributed to multi-agent systems — "specialised roles", "a team of experts", "emergent collaboration" — is not a mechanism. The model is the same model. A prompt saying "you are a security expert" does not add capability; it biases what the output attends to, which is occasionally useful and not what the marketing claims.
What a second agent costs
- It starts empty. Everything it needs must be in its instructions, written by whoever spawned it, usually more briefly than you would.
- It cannot see its siblings. Four agents solving four parts of one problem produce four locally reasonable, mutually inconsistent solutions.
- The summary is lossy. What comes back to the parent is a secondary source, with everything that implies.
- Review does not parallelise. This is the binding constraint and the one people discover last.
The decision procedure
1. Would this work pollute my context with material I will not need?
yes -> subagent. This is the strongest case.
2. Does this need judgement that must be independent of how the work
was done?
yes -> separate session. Review, critique, adversarial checking.
3. Are these pieces genuinely independent - different files, no shared
decisions, no need to agree on naming?
yes -> parallel agents may help
no -> sequential. You will spend more reconciling than you saved.
4. Can I review all the output in the time I have?
no -> do not start more runs. This overrides everything above.Most tasks stop at 1 or 2, which is fine — those are where the value is.
The consistency problem, concretely
| Task | Parallel? | Why |
|---|---|---|
| Migrate 23 files to a new API, exemplar given | Yes | Each file is independent; the exemplar makes the decisions |
| Write tests for 8 existing modules | Yes | Independent, and the modules already exist to constrain them |
| Build 4 new endpoints for one feature | No | They must agree on error shape, naming, validation, auth |
| Explore 3 areas of an unfamiliar codebase | Yes | Read-only, and only summaries come back |
| Design a schema and build against it | No | Sequential by nature; one output feeds the next |
The pattern: parallel works when the decisions are already made — by an exemplar, by existing code, or by a spec precise enough to leave nothing open. Where decisions remain, parallelism multiplies them.
The honest summary
Subagents for context isolation: adopt immediately, high value, low risk. Separate sessions for review: adopt immediately, highest value in the module. Parallel implementation: useful for mechanical work with a strong exemplar, bounded by your review capacity. Elaborate multi-agent architectures with roles and negotiation: almost always slower, more expensive and less correct than one well-specified sequential chain, and the rest of this module is about the narrow cases where that is not true.
Exercise
Take the last three things you used multiple agents for — or considered doing. Run each through the four-step decision procedure honestly.
For any that fail step 3, write down what decision the pieces would have had to agree on. That decision is the thing you would have spent your evening reconciling.
Worked solution
Three real cases, run through the procedure.
Step 1: no, each agent needs its own context anyway.
Step 2: no.
Step 3: FAILS. The pages must agree on: form layout, validation timing
(on blur or on submit), how errors render, the save-and-toast
pattern, and where shared state lives.
Ran it anyway, as an experiment. Five agents, five worktrees, 20 minutes.
Result: five working pages with four different validation approaches,
three different toast implementations and two new shared hooks that did
the same thing with different names.
Reconciliation took two and a half hours - longer than building them
sequentially would have taken.Step 1: yes, and step 3 passes - the modules exist, so behaviour is
already determined; nothing has to be agreed.
Ran 4 at a time, each with the same instruction and the same exemplar
test file.
Result: consistent, and the whole batch took 35 minutes against a day
of my time. The only reconciliation was two duplicate test helpers.Step 1: YES - this is the strongest case. Each exploration would have
cost 30-50k tokens in my window and I needed none of it after
the answer.
Three subagents, read-only, each returning under 20 lines.
My session went from a projected 120k to about 9k, and I could then
do the actual design work in a sharp window.The instructive comparison is cases 1 and 2, which look superficially identical — N similar things, done at once. The difference is that in case 2 the answer was already determined by existing code, and in case 1 every page was making the same five decisions independently.
The retrospective fix for case 1 is not "do not parallelise". It is: build one page first, sequentially, and then parallelise the other four against it as an exemplar. That turns the open decisions into settled ones and moves the task from failing step 3 to passing it. Doing exactly that later took 25 minutes plus the original page, with no reconciliation at all.
Takeaways
- A second agent buys a separate context window and independence from your reasoning. Nothing else.
- Parallel work succeeds only when the decisions are already made by an exemplar, existing code, or an airtight spec.
- Build one instance sequentially, then parallelise the rest against it as an exemplar.
- Your review capacity overrides every other consideration.
Check yourself
Why does building four new endpoints in parallel usually cost more than it saves?
Each agent makes the same open decisions independently and reasonably, producing four defensible but incompatible answers. Reconciling them costs more than sequential work would have. Settle the decisions first — usually by building one — and parallelism becomes safe.
A course by Pieter Zandbergen