Instrument your setup
You cannot improve a process you cannot see. Three cheap instruments worth wiring up today.
Most people work with agents entirely by feel — which is why opinions about what works vary so wildly. Three small instruments turn feel into evidence.
1. Token counts per session
Find your harness's usage display and glance at it at the end of every session for a week. Record three numbers in your log: total input, total output, and turn count.
What you are looking for is the outlier. The session that cost five times the others is the interesting one, and the cause is almost always identifiable: a directory dump, a verbose test run, a debugging detour you never cleared.
2. Session transcripts
Most harnesses can save or export a transcript. Keep them for a couple of weeks. They are the only way to answer "what did I actually ask" after the fact, and the only honest input to your instruction files.
2026-03-04 feat/backoff
turns: 6 in: 240k out: 9k result: accepted
corrections:
- had to say "use the existing Result type, not exceptions" -> GENERAL
- had to say "tests go in __tests__, not next to source" -> GENERAL
- had to say "the retry cap is 5, see config/queue.ts" -> specific
note: it invented a withRetry() helper that does not exist. caught by tsc.The GENERAL tags are the point. After a week you will have ten to twenty of them and they write your project brief for you — from evidence, not from imagination. That is the input to Lesson 5.1.
3. A correction counter
The crudest and most useful metric: how many times per session do you have to correct the agent on something that is not task-specific? Track it daily. It is the single best measure of whether your steering is working, and it should trend down as Module 05 lands. If it does not, your instruction file is being written but not read — a real and common failure worth catching early.
Watch out
Do not turn this into a dashboard project. Three numbers and a bullet list in a markdown file. The moment instrumentation becomes work, it stops happening, and a log you keep for two weeks beats a system you abandon after two days.
What good looks like after a month
- Median session under 50k tokens and under eight turns.
- General corrections down to one or two a day, because the rest are encoded.
- You can name your own dumb-zone threshold from experience rather than from this course.
- Your commit history shows one session per commit, and each diff is reviewable in under five minutes.
Try it
Create notes/agent-log.md with the entry format above and commit to filling it for ten sessions. Module 05 opens by asking you to read it back — and it is much better with real data in it.
Takeaways
- Track tokens, turns, and general corrections — three numbers, one file.
- Tag corrections GENERAL or specific; the general ones become your project brief.
- Keep instrumentation small enough that you actually keep doing it.
Your general-correction count is not falling even though you wrote an instruction file. What are the two likely causes?
Either the file is not being loaded — wrong filename, wrong location, or a harness that does not read it — or it is loaded but too long and vague to act on, so it is present in context without changing behaviour. Check loading first: it is binary and takes thirty seconds to test.
Module 03 of this course ends here
The Next Steps goes under it: building your own tools: check commands as products, custom commands, skills that survive, hooks, and small MCP servers.
See what is in it →A course by Pieter Zandbergen