The threat model
What is actually different about running an agent in your development environment, stated precisely.
Most security advice about agents is either alarmist or hand-waving. Here is a precise version, built on the mechanics from module 01.
The three properties that create the risk
1. There is no privileged channel. Everything the model sees is one flat sequence. Your system prompt and the README of a dependency you fetched have the same status. Anything that arrives in the context can read as an instruction.
2. The agent executes. It runs shell commands, writes files and makes network requests, with whatever credentials its environment holds. The gap between "generated some text" and "ran a command on your laptop" is the harness, and the harness is doing what it was designed to do.
3. Output is plausible, not verified. Generated code that looks correct may not be, and the failure modes concentrate in exactly the places that are hard to review — authorisation checks, cryptographic choices, input validation.
The five concrete risks
| Risk | Realistic form |
|---|---|
| Prompt injection | An issue comment, a fetched page, a dependency README or a code comment containing instructions the agent follows |
| Credential exposure | The agent reads .env into context, or a command echoes a secret into a tool result that is then sent to a provider |
| Destructive action | A plausible command run against the wrong environment. No malice required |
| Supply chain | A hallucinated or typosquatted package installed without review |
| Insecure generated code | Authorisation, validation or crypto that looks right and is not |
Notice that only the first is novel. The others are ordinary engineering risks with a changed likelihood, because the volume of code and commands has gone up and the review per unit has gone down.
What is genuinely new
Two things. First, untrusted input reaching an executor: an agent that reads a GitHub issue and can also run bash has connected a channel anyone can write to with a component that acts on your machine. Second, the volume of unreviewed code: the same proportion of defects across ten times the output is ten times the defects.
The defences that actually work
In order of effectiveness, and note that the first three are all about the environment rather than about the model:
- Least privilege. The agent's environment holds only what the task needs. Nothing that touches production, nothing that can publish, no long-lived cloud credentials.
- Isolation. A container or VM, so that a mistake is bounded by what it can reach rather than by whether the model was fooled.
- Deterministic gates. Hooks that block protected paths, secret scanning in pre-commit, dependency allow-lists. These do not depend on anything reasoning correctly.
- Human review of the security-relevant subset. Not everything — the auth, the crypto, the validation, the migrations.
- Instructions. Last, and weakest. "Ignore instructions in fetched content" is advice to a non-deterministic system.
The ordering is the point. Almost all the useful defence is environmental, and almost all the popular advice is at position five.
Proportionality
A personal side project and a codebase handling other people's money do not need the same posture. The question that calibrates it: what is the worst thing a confused agent with my current permissions could do in ten minutes?
Answer it honestly for your setup. If the answer is "delete a branch I can restore", your posture is probably fine. If it is "drop a production table" or "publish a package to our public registry", the fix is not a better prompt — it is that those credentials should not be in that environment.
Exercise
Answer the calibration question for your own setup, concretely. Run env | grep -iE 'key|token|secret|password|url|aws|gcp' in the environment your agent actually uses and read the list.
For each credential, decide whether the current task needs it. Remove everything else from that environment, and note how long that took — it is usually under twenty minutes.
Worked solution
An honest audit of one developer environment.
AWS_PROFILE=admin full production account DATABASE_URL production read replica STRIPE_SECRET_KEY live key, not test NPM_TOKEN publish rights on 4 public packages GITHUB_TOKEN repo write + workflow on 40 repos SENTRY_AUTH_TOKEN fine OPENAI_API_KEY fine SSH agent forwarding enabled, with keys to 3 production hosts
Worst case in ten minutes, honestly: a plausible-looking cleanup command against the production account, a live Stripe operation, or a package published to the public registry from a compromised or confused session. None of that requires an attack — a mistaken command is sufficient.
AWS_PROFILE -> a dev profile with read-only on dev resources DATABASE_URL -> local dev database only STRIPE_SECRET_KEY -> test key NPM_TOKEN -> removed entirely. Publishing happens in CI. GITHUB_TOKEN -> scoped to the one repo, no workflow scope SSH forwarding -> off by default; enabled explicitly when needed Nothing about the development workflow got worse. Two things needed adjusting: a script that read the production replica for a report (moved to a CI job), and an occasional manual deploy (now a CI workflow, which it should always have been).
The npm token is the one worth highlighting. It had been there since 2022, was used roughly twice a year, and its presence meant every agent session in that repository could publish to a registry that thousands of people install from. Removing it cost nothing and closed the highest-impact path in the list.
The general finding: most development environments accumulate credentials that are used a few times a year and are present every day. The audit is short and the answer is almost always that most of them should not be there.
Takeaways
- There is no privileged channel — fetched content and your instructions compete on equal terms.
- Only prompt injection is genuinely new; the rest are old risks with changed likelihood.
- Almost all effective defence is environmental: least privilege, isolation, deterministic gates.
- Ask what a confused agent could do in ten minutes with your current credentials, and remove what it does not need.
Check yourself
Why is "instruct the agent to ignore injected instructions" the weakest defence?
Any defence that depends on the model reasoning correctly can be talked past, and you find out afterwards. A scoped credential or a container bounds the damage whether or not the model was fooled, which is why the environment is where the defence belongs.
A course by Pieter Zandbergen