Agent Engineering
Module 05 · Steering/Lesson 5.8/3 min

Steering anti-patterns

Six ways briefs fail, and how to tell whether yours is actually working.

The six

  1. The novel. Four hundred lines covering everything. Costs attention in every session and gets skimmed. Symptom: you cannot remember what is in it. Fix: the four-question filter from Lesson 5.2, ruthlessly.
  2. The aspiration. "Write clean, well-tested code." Unfalsifiable, so it changes nothing. Fix: replace with a rule a check could verify.
  3. The lie. Rules the codebase does not actually follow. The agent reads the brief, reads the code, finds a conflict, and resolves it unpredictably — often by copying the code. Fix: either make the code match, or document the exception explicitly under known rough edges.
  4. The stale pointer. Paths and commands that no longer exist. Costs a failed tool call and erodes the credibility of everything around it. Fix: put a command in the brief that would fail if it were stale, and run it occasionally.
  5. The personality prompt. "You are a 10x engineer who never makes mistakes." No measurable effect on current models, and it displaces something useful. Fix: delete.
  6. The private brief. Lives in one person’s settings, so two teammates get different results from the same task and new joiners get none of it. Fix: check it in.

Testing whether the brief works

Three tests, in increasing strength.

1. is it loaded?
Fresh session:
  "Without reading any file, quote the rule in your instructions about
   database migrations."
Cannot quote it -> it is not being loaded. Fix the filename or location first;
everything else is wasted until this passes.
2. does it change behaviour?
Give the agent a task that would violate a specific rule if it were ignored.
Watch whether it follows the rule unprompted.
A rule that is loaded but not followed is usually too vague, or buried in
too much text.
3. does it survive a long session?
Run a long, messy session and give the same task at the end.
Rules that fail here belong on the check ladder (Lesson 5.5), not in prose.

The measurement that matters

Your general-correction count from Lesson 3.7. It should fall week over week as you encode more. If it plateaus, you have reached the limit of what prose can do, and the remaining corrections need checks or better task scoping instead.

Watch out

Do not add a brief line in response to a single incident. One-off mistakes are non-determinism (Lesson 2.6), and reacting to each one is how you end up with the novel. Wait until you have corrected the same thing twice.

Try it

Run all three tests on your brief right now. Test 1 takes thirty seconds and fails more often than people expect — wrong filename, wrong directory, or a harness that needs it enabled.

Takeaways

  • Long, aspirational, untrue, stale, flattering, or private — the six ways briefs fail.
  • Test loading first: a brief that is not loaded makes every other question moot.
  • Rules that fail late in long sessions belong in checks, not prose.
  • Add a line only after the second identical correction.
Why is a brief that contradicts the codebase worse than one that omits the rule entirely?

Because it creates two conflicting sources and the resolution is unpredictable — the agent may follow the brief, follow the code, or split the difference, differently each session. Omission at least leaves one consistent source. State the exception explicitly instead.

Module 05 of this course ends here

The Next Steps goes under it: orchestration and multi-agent work, the check ladder, property testing, and a personal eval suite you can run against any new model.

See what is in it →

A course by Pieter Zandbergen