Why debugging is different
Every other task starts from a specification. Debugging starts from a contradiction, and that changes the whole method.
Building has a specification: you know what you want and the question is how. Debugging has a contradiction: the system is doing something that should be impossible given your understanding of it, and the real task is finding out which part of your understanding is wrong.
That difference matters because an agent's core competence — producing a plausible continuation — is exactly the wrong instinct here. The plausible explanation is usually the one you already considered and dismissed.
The three ways agent-assisted debugging fails
1. Pattern matching on the symptom. "Undefined is not a function" produces the five most common causes of that message in general, none of which is your cause. The agent has not reproduced anything; it is completing a sentence.
2. The speculative-fix spiral. Change something, see if it helps, change something else. Forty minutes later there are six unexplained edits, the original bug is still there, and the working tree is unrecoverable.
3. Confident wrong causes. The agent identifies a cause, explains it fluently, proposes a fix, the fix appears to work, and the real bug resurfaces next week — because the change perturbed timing rather than fixing anything.
What agents are genuinely good at
The failures above are all about concluding. The things agents do well in debugging are all about gathering:
- Generating hypotheses. Producing five plausible causes ranked by likelihood is a real strength.
- Exhaustive reading. Every caller, every path, every place a value could be set.
- Building instrumentation. Writing the logging, the assertion, the reducer over the log file.
- Reading history. Bisecting, blaming, finding when behaviour changed.
- Reading large evidence. A thousand-line trace, a heap dump summary, a flame graph description.
So the method is: let it generate and gather; you decide. Everything in this module is a way of enforcing that split.
The one rule
No change to the code until a hypothesis has been confirmed by evidence.
Not "until a hypothesis exists" — hypotheses are cheap and an agent will produce them indefinitely. Until one has been confirmed, by an observation that would have come out differently if it were false.
This single rule eliminates the speculative-fix spiral entirely, and it is the reason the four-step structure in the next lesson works.
The cost of getting it wrong
Debugging is where sessions go worst, and the mechanism is the one from module 02: a debugging session accumulates stack traces, ruled-out hypotheses, diagnostic output and abandoned edits faster than any other kind of work. By turn twenty the window is full of material that is actively misleading, and the agent's suggestions get worse precisely as the problem gets harder.
Which gives the second rule, and it is the same one as the free course: debugging happens in its own session, and ends with knowledge rather than a fix.
Exercise
Look back at the last bug that took you more than an hour with an agent. Count the code changes made before a cause was confirmed by evidence.
Then ask the harder question: what was the first observation that would have distinguished between your top two hypotheses, and how long before you made it?
Worked solution
A real post-mortem on a bug that took three hours: intermittent 500s on one endpoint, roughly one request in two hundred.
0:00 described the symptom to the agent
0:03 it proposed a race in the connection pool. Plausible.
0:06 changed the pool size. Deployed to staging. Waited.
0:25 still happening. Proposed a timeout issue.
0:28 changed the timeout. Waited.
0:50 still happening. Proposed retry storm.
0:55 added a circuit breaker (60 lines). Waited.
1:20 still happening.
1:25 proposed it might be the load balancer.
...
2:10 I gave up on the agent and added one log line recording the
request id, the pool state and the upstream response code.
2:35 the log showed it was always the same upstream, always returning
499, always after exactly 30s.
2:40 cause found: an upstream timeout, not a pool problem at all.
3:00 fixed.Code changes before the cause was confirmed: three, totalling about ninety lines, all reverted. Time from the first observation that discriminated between hypotheses to the fix: thirty minutes. Time before that observation was made: two hours and ten minutes.
A similar bug two weeks later:
0:00 described the symptom, with the instruction: "list 3 hypotheses
with evidence for each, and for the top one, tell me the cheapest
observation that would confirm or rule it out. Do not edit
anything."
0:04 three hypotheses. The discriminating observation for the top one
was a single log line.
0:09 added the log line, deployed.
0:31 log showed the top hypothesis was WRONG, and the shape of the
data pointed at the second.
0:38 confirmed the second with one more observation.
0:55 fixed.Fifty-five minutes against three hours, and the interesting number is that the first run's top hypothesis was wrong too — the difference is entirely in how long it took to find that out. Two hours of speculative fixes versus twenty minutes of one log line.
Takeaways
- Debugging starts from a contradiction, not a specification — the plausible answer is usually already dismissed.
- Let the agent generate and gather; you decide. It is strong at hypotheses and evidence, weak at concluding.
- No code change until a hypothesis is confirmed by an observation that could have come out otherwise.
- Debug in a dedicated session that ends with knowledge, not a fix.
Check yourself
Why is "no code change until a hypothesis is confirmed" a stronger rule than "form a hypothesis first"?
Requiring a hypothesis costs nothing — there is always one available, and changing code to test it is the reflex. Requiring a confirming observation forces the cheap, informative step that actually distinguishes between candidates, which is the step that gets skipped.
A course by Pieter Zandbergen