Next Steps
Module 07 · Legacy and Large Codebases /Lesson 7.1 /5 min

What legacy actually changes

Every assumption in the previous six modules quietly depends on things a legacy codebase does not have.

The techniques so far assume a fast check command, decent test coverage, clear module boundaries and someone who knows why the code is the way it is. A genuinely legacy system has none of those, and the honest response is not to apply the same techniques harder.

The four missing foundations

1. No feedback loop. No tests, or tests that take forty minutes, or tests that fail for reasons nobody understands. Without it the agent cannot self-correct, and every error surfaces in your review instead.

2. No reliable context. Comments describe behaviour that changed in 2019. The README describes an architecture that was replaced. The names lie: getUser writes to three tables and sends an email. An agent reading this codebase is reading confident, specific, wrong secondary sources everywhere it looks.

3. Behaviour is load-bearing and undocumented. Some of what looks like a bug is relied upon by a customer, a downstream report, or a script someone runs on Fridays. "Fixing" it breaks production in a way no test catches.

4. Nobody knows why. The person who wrote it has left. The commit message says "fix". The odd-looking condition on line 340 is either dead code or the thing that stops a specific customer's data corrupting, and there is no way to tell from the code.

What this implies

In a healthy codebaseIn a legacy one
Explore, then implementExplore, verify what you learned, then implement
Tests tell you if you broke itNothing tells you. Build the net first.
Read the code to understand intentRead the code, the history, and the issue tracker
Refactor freely with a green suiteCharacterise first; refactor only behind the net
Odd code is probably wrongOdd code is probably load-bearing until proven otherwise

The last row is the attitude shift that matters most. In a healthy codebase, code that looks wrong usually is. In a legacy one, the default assumption must be inverted — and an agent, whose training prior is full of clean code, will reliably assume the opposite.

The instruction that prevents the most damage

AGENTS.md for a legacy area
## src/legacy/
This code predates our current conventions and has no meaningful test
coverage. Assume every oddity is deliberate until proven otherwise.

- Do not "modernise", reformat, rename, or tidy anything here.
- Do not fix bugs you notice unless they are the task. Report them.
- Do not trust comments or names. Verify behaviour by reading the code
  or by running it.
- Any change must be accompanied by a characterisation test written
  BEFORE the change, capturing current behaviour including anything
  that looks wrong.

Without that block, the single most common legacy failure is an agent making a 40-line fix and a 400-line improvement in the same diff — at which point you cannot review either.

Where agents are unusually good here

It is not all worse. Legacy work has a large component of tedious, mechanical reading that humans avoid and agents do not mind:

  • Tracing a path through code nobody understands. Tireless, and forced to quote.
  • Finding every caller of something before you change it. Exhaustive in a way manual search is not.
  • Writing characterisation tests — genuinely dull, genuinely valuable.
  • Reading git history to find out why a line exists.
  • Mechanical migration once you have established the pattern.

The pattern for the rest of this module: use agents heavily for the reading and the net-building, cautiously for the changing, and never for deciding what is safe to touch.

Exercise

Pick the area of your codebase you are most reluctant to change. Write the legacy instruction block for it, adapted to what is actually true there.

Then ask an agent, with that block loaded, to make a small change in it. Note whether the block stopped the improvement reflex — and if it did not, which clause was missing.

Worked solution

Area: a 2,800-line invoice generator written in 2018, no tests, three known bugs that finance depends on.

without the block
Task: "add a VAT-exempt line item type to the invoice generator."

Diff: 41 lines for the feature, 380 lines of other changes.
  - reformatted the whole file to the current prettier config
  - renamed 11 variables ("d" -> "documentDate" etc.)
  - replaced a hand-rolled currency rounding helper with our Money type
  - "fixed" a condition that excluded zero-amount lines from the total

Every change defensible. The last one would have altered every invoice
we generate - finance relies on zero-amount lines being listed but not
summed, which is why that condition exists. There was no comment.
with the block
Same task. Diff: 44 lines, all feature.

Plus, at the end of the session, unprompted:

"Three things in this file look wrong but I have not changed them, per
 the instructions:
   - line 412 excludes zero-amount lines from the total but includes
     them in the rendered list
   - line 588 rounds half-down, which differs from src/lib/money.ts
   - line 1204 catches and ignores an error from the PDF renderer
 Reporting rather than fixing."

All three are real, and all three are load-bearing. Zero-amount exclusion is deliberate. Round-half-down matches what the tax authority expects for this jurisdiction and differs from the rest of the system for a reason nobody had written down. The swallowed PDF error is the only genuine defect, and it is now a ticket.

The clause that did the work was not "do not modernise" — it was "report them". Without somewhere for the improvement reflex to go, the agent either acts on it or drops it. Giving it a legitimate outlet produced a better inventory of that file's oddities than four years of people working on it had managed.

Takeaways

  • Legacy code invalidates the assumptions the previous modules rest on — build the foundations before using the techniques.
  • Invert the default: odd code is load-bearing until proven otherwise.
  • Comments and names in legacy code are confident, specific, unreliable secondary sources.
  • Give the improvement reflex an outlet: "report, do not fix" produces a valuable inventory.

Check yourself

Why is "report them, do not fix them" more effective than "do not fix them"?

A course by