Next Steps
Module 03 · Building Your Own Tooling /Lesson 3.1 /5 min

Why you should build tools

The gap between a good agent setup and an average one is mostly tooling, and it is cheaper to build than it has ever been.

Most people accept whatever tools their harness ships with and then spend months working around the gaps. That made sense when a bespoke tool cost a day. It costs twenty minutes now, and the agent writes most of it.

The three things tooling buys

1. Determinism where it matters. A tool does the same thing every time. Every piece of work you move from "the agent figures it out" to "the agent calls this" becomes repeatable and reviewable. The migration script that runs the same five steps in the same order does not have a bad day.

2. Context compression. A well-designed tool replaces a long exploratory sequence with one call and a small result. The difference between an agent reading four files to work out your deployment state and calling deploy-status is 18,000 tokens against 200 — and 18,000 tokens of attention budget preserved for the actual work.

3. Encoded judgement. A tool that refuses to do something, or returns an error saying what to do instead, encodes your standards in a form that cannot be forgotten in the dumb zone or talked past in the twentieth turn.

What to build, in order of return

BuildReplacesEffort
A fast, quiet check commandThe agent guessing whether it broke somethingAn hour, once
Project-specific lint rulesCorrecting the same convention forever20 min each
A repo status toolSix exploratory commands every session30 min
Custom commands for your ritualsRetyping the same opening message10 min each
Skills for multi-step proceduresThe procedure being done wrong every third time30 min each
A small MCP serverPasting data from another system by handHalf a day

The ordering is deliberate. The first two are where nearly all the value is, and both are unglamorous. People reach for the MCP server first because it feels like engineering; the check command is worth more.

The test for whether something should be a tool

Three questions. Two yeses is enough.

  1. Has the agent done this by hand more than three times? Repetition is the signal, and you will have it in your session log.
  2. Does doing it by hand cost more than 2,000 tokens? Then a tool pays for itself in context alone.
  3. Does it go wrong in a way a tool could prevent? Then the tool is buying correctness, not just speed.

The build loop

The agent builds its own tools, which feels circular and is not — the tool is deterministic once written, however it was written.

a twenty-minute tool
"Write a script bin/repo-status that prints, in under 30 lines:
   - current branch and whether it is behind origin
   - uncommitted file count
   - whether the check command currently passes (just pass/fail)
   - the 5 most recently modified source files
 Plain text, no colour, no progress output. Fail with a clear message
 if run outside the repository root."

then add to AGENTS.md:
  Run bin/repo-status at the start of any session that will change code.

That single tool replaces a sequence most agents perform badly at the start of every session, and it costs about 150 tokens to run.

The cost side

Tooling is not free. Every tool has a definition in the system prompt of every session (module 01, lesson 8), needs maintaining, and can go stale. Two rules keep this in check:

  • Tools live in the repository, in version control, so they are reviewed and they travel with the project.
  • Delete tools you have not used in a month. An unused tool is pure cost, charged on every request.

Exercise

Read back two weeks of your session log and find the sequence of commands your agent performs most often at the start of a session. Have it write that as a single script.

Then measure: tokens for the old sequence, tokens for the tool, and how many sessions a week you run. Multiply. That is the annual return on twenty minutes.

Worked solution

The sequence found in one real log, performed at the start of 19 of 24 sessions:

before
git status                      ~300 tokens  (often 40+ untracked files)
git log --oneline -10           ~250
ls -la src/                     ~400
cat package.json                ~900
npm test                     ~3,200  (full output, to see if things pass)
                              -----
                              ~5,050 tokens, 5 turns, every session
bin/repo-status, written by the agent in one turn
#!/usr/bin/env bash
set -euo pipefail
cd "$(git rev-parse --show-toplevel)"
echo "branch:    $(git branch --show-current) ($(git rev-list --count @{u}..HEAD 2>/dev/null || echo 0) ahead, $(git rev-list --count HEAD..@{u} 2>/dev/null || echo 0) behind)"
echo "dirty:     $(git status --porcelain | wc -l) files"
echo "check:     $(npm run --silent check >/dev/null 2>&1 && echo PASS || echo FAIL)"
echo "scripts:   $(node -p "Object.keys(require('./package.json').scripts).join(', ')")"
echo "recent:"
git log --format='  %h %s' -5
echo "touched:"
git diff --name-only HEAD~5 2>/dev/null | head -8 | sed 's/^/  /'
after
bin/repo-status                 ~180 tokens, 1 turn

The arithmetic: 4,870 tokens saved × 19 sessions a week × 46 working weeks is roughly 4.2 million input tokens a year, for twenty minutes of work. The cost saving is pleasant; the attention saving is the real return, because every one of those sessions now starts with 5k more of its budget intact.

One design note worth copying: the script prints check: PASS rather than the test output. The agent almost never needs the output — it needs to know whether it is starting from green, so that a failure later is attributable to its own change. Printing the answer instead of the evidence is the whole trick of agent-facing tool design.

Takeaways

  • Tooling buys determinism, context compression and encoded judgement — in that order of importance.
  • The check command and project-specific lint rules are worth more than anything glamorous.
  • Build a tool when the agent has done it by hand three times, or when doing it costs more than 2,000 tokens.
  • Print the answer, not the evidence.

Check yourself

Why should a status tool print "check: PASS" rather than the test output?

A course by