Running it where it cannot hurt you
Worktrees, containers, and throwaway environments — how to raise autonomy without raising risk.
Lesson 1.3 introduced the two dials: autonomy and blast radius. This lesson is about the second one, because lowering it is what makes raising the first one sane.
Level 1: git worktrees
The cheapest isolation available, and badly underused. A worktree gives you a second checkout of the same repository on a different branch, in a different directory. The agent works there; your main checkout is untouched and stays usable.
git worktree add ../myapp-agent-backoff -b feat/backoff cd ../myapp-agent-backoff # run the agent here. your main checkout keeps working. # when done: git worktree remove ../myapp-agent-backoff
This also unlocks parallel sessions: two agents on two unrelated tasks in two worktrees, neither stepping on the other's files. That is a real throughput gain and it costs nothing.
Level 2: a container
A sandbox — devcontainer, plain Docker, a VM — bounds what any approved command can reach. The properties that matter:
- Only the repository is mounted. Not your home directory, not your SSH keys.
- Only the credentials the task needs, and none that touch production.
- Network egress restricted to what the build actually requires, if you can manage it.
- Disposable. If it gets into a bad state, you delete it.
Inside a container like that, broad auto-approval is a reasonable engineering choice rather than a gamble. Outside one, it is a gamble regardless of how good the model is.
Level 3: a clean remote environment
Hosted and background agents run in a fresh environment per task and hand you a pull request. Maximum isolation, and it forces good specs, since you cannot course-correct mid-run. The trade is latency and the loss of the tight interactive loop.
The credential question
Whatever level you use, take five minutes on this: what is in the agent's environment? Cloud credentials, a production database URL, a package-registry publish token, an SSH agent with forwarded keys. Agents run shell commands on your behalf, and a plausible-looking command that touches the wrong environment does not need malice to cause an incident.
Watch out
Prompt injection is a live risk once agents read external content — issues, web pages, dependency READMEs, code review comments from strangers. Instructions inside fetched content can and do attempt to steer the agent. Isolation is the defence that does not depend on the model resisting the attempt.
Try it
Set up a worktree and run one task in it. Then run env | grep -iE 'key|token|secret|password|url' in the environment your agent actually uses, and decide honestly whether everything in that list should be reachable by a generated shell command.
Takeaways
- Worktrees are near-free isolation and enable parallel sessions.
- A container with scoped credentials is what makes broad auto-approval defensible.
- Audit what is in the agent’s environment; isolation beats trusting the model.
Why is isolation a better defence against prompt injection than instructing the agent to ignore untrusted instructions?
Because the instruction is advice to a non-deterministic system and the isolation is a property of the environment. If the agent is fooled, a scoped container limits what the mistake can reach; a rule in a prompt limits nothing once it has been talked past.
A course by Pieter Zandbergen