Agent Engineering
Module 04 · Fundamentals/Lesson 4.6/3 min

Working with tests

Agents write tests eagerly and badly by default. Two rules fix most of it.

Test generation looks like the perfect agent task and mostly is not, for one structural reason: a test written after the implementation, by the thing that wrote the implementation, tends to assert what the code does rather than what it should do. If the code has a bug, the test enshrines it. You get coverage numbers and no safety.

Rule one: specify the cases, delegate the writing

You decide what must be true. The agent writes the mechanics.

prompt
Write tests for parseRow() covering exactly these cases:
  1. ISO date -> parsed correctly
  2. DD/MM/YYYY -> parsed correctly
  3. ambiguous 01/02/2026 -> uses the configured locale, not a guess
  4. malformed date -> returns an error, does not throw
  5. empty date column -> row still imports, date is null

Assert on behaviour and return values only. Do not assert on internal calls
or on how the function is implemented. No mocking of the module under test.

The last two lines prevent the most common failure: tests that mock so thoroughly they verify the mocks.

Rule two: make the test fail first

The one habit that separates real tests from decorative ones.

prompt
Before implementing, write the test and run it. Show me it failing, with the
failure message. Then implement, and show me it passing.

A test that has never failed is a test you have no evidence about. This is ordinary test-driven discipline, and it matters more with agents than without, because the volume of generated tests makes each one less likely to get individual scrutiny.

When you inherit untested code

Backfilling tests is a legitimately good agent task, with one ordering constraint: characterise before you change. Have the agent write tests that document current behaviour, review them yourself for anything that looks wrong-but-intended, and only then refactor. Now the test suite is a safety net rather than a record of the bug.

prompt
Write characterisation tests for this module: capture what it currently does,
including behaviour that looks like a bug. Add a comment marked SUSPECT on any
assertion where the current behaviour seems wrong, and list those separately.
Do not change any source file.

That SUSPECT list is often the most valuable output of the session — it is a bug inventory you did not have to find by hand.

Watch out

Never let the same session both fix a failing test and edit the code under test without you reading both halves of the diff. The path of least resistance is to change the assertion. Read the test diff first, every time, as in Lesson 3.3.

Try it

Find a module in your repo with poor coverage. Run the characterisation prompt. Review the SUSPECT list — most people find at least one real bug they had been shipping for months.

Takeaways

  • You specify the cases; the agent writes the mechanics.
  • Demand a failing run before the implementation — an untested test is not evidence.
  • Characterise legacy behaviour before refactoring, and ask for a SUSPECT list.
What is wrong with a test generated from a finished implementation?

It is derived from the code rather than from the requirement, so it asserts whatever the code happens to do, bugs included. It then blocks refactoring while proving nothing about correctness — the worst of both properties. Specify the cases from the requirement instead.

A course by Pieter Zandbergen