Agent Engineering
Module 06 · Shipping/Lesson 6.6/3 min

Unattended runs

Letting an agent work while you do something else — and getting back something you can actually use.

AFK work pays off when three things are true: the task is well specified, the environment is contained, and the result is verifiable without you. Missing any one of them turns an unattended run into a long wait for a diff you throw away.

The preflight

before you walk away
[ ] The ticket is complete: a cold agent could execute it with no questions.
[ ] The check command passes right now, before any changes.
[ ] Scope is bounded: named files, explicit "do not" list.
[ ] The environment is contained: worktree or container, no production creds.
[ ] There is an escape hatch: "if X, stop and report" for the likely blockers.
[ ] The output is a diff you can review, not a merged branch.

The escape hatch is the line people forget. An agent that hits an unexpected blocker will work around it, and a workaround produced without supervision is exactly the class of change you least want to inherit.

prompt
Work through specs/scheduled-exports/t4-monthly-rules.md.
Run 'npm run check' after each logical step and keep it green.
Commit at each green step with a descriptive message.

Stop and write a note in NOTES.md instead of continuing if:
  - the ticket turns out to be ambiguous
  - you need to change anything outside src/export/
  - you need a new dependency
  - the check command fails three times on the same issue

Do not merge, do not push, do not open a PR.

Commit as you go

Instruct the agent to commit at each green step. This gives you a reviewable history instead of one enormous diff, it lets you accept the first four commits and drop the fifth, and it acts as a progress log if the run ends badly.

Reviewing a cold result

You are reviewing without having watched it happen, so read in this order:

  1. NOTES.md — did it stop, and why?
  2. The commit list — does the shape of the work match the ticket?
  3. The diff stat — is anything outside the scoped files?
  4. The tests — do they assert what the ticket demanded, or what the code does?
  5. The implementation, last.

Then run a fresh-session automated review over the whole diff (Lesson 3.3) before your own detailed read. It is cheap and it catches the mechanical problems, so your attention goes to judgement calls.

Parallel unattended work

Several independent tickets can run at once in separate worktrees. The constraint is not the agent — it is you. Three runs finishing together produce three diffs to review, and review does not parallelise. Start as many as you can review in the time you have.

Watch out

Never point an unattended run at a task whose failure mode is silent. Data migrations, anything touching money, anything that writes to a system you cannot roll back. If the checks cannot catch a wrong result, being present is not optional — it is the only control you have.

Try it

Run one ticket unattended with the full preflight. Afterwards, classify every problem you find as a spec gap, a scope failure, or a genuine model error. Most people find the first category dominates — which tells you exactly what to improve.

Takeaways

  • Unattended work needs a complete ticket, a contained environment, and verifiable output.
  • Always give explicit stop-and-report conditions.
  • Commit at each green step so the result is reviewable in pieces.
  • Start only as many parallel runs as you can review.
Why does parallel unattended work stop scaling so quickly?

Because review does not parallelise. The agent time is cheap and concurrent; your attention is serial and is the actual constraint. Three simultaneous runs produce three cold diffs, and the third gets reviewed worst.

A course by Pieter Zandbergen