Non-determinism and effort
Why the same prompt gives different answers, and what the reasoning dial actually buys you.
Run the same prompt twice against the same model and you can get different output. This is non-determinism, and it has two sources: sampling during generation, and variation in how the provider serves the request (batching, hardware, routing). Even at temperature zero, providers do not generally guarantee bit-identical results.
What this means for your process
- A success is not a guarantee. A prompt that worked once will not always work. If you need reliability, the reliability has to come from the check, not the prompt.
- A failure is not a verdict. Before concluding the agent cannot do something, retry once in a fresh session. Surprisingly often the second attempt is fine, and you have learned that the task sits near the edge of what it does reliably — useful information.
- Re-rolling is a legitimate technique, but only for cheap, verifiable tasks. Re-rolling a 40-minute refactor you then have to review three times is not a technique, it is a lottery.
Effort
Effort is a dial exposed by most current harnesses (sometimes called reasoning effort, thinking budget, or extended thinking). Turned up, the model generates a larger internal reasoning pass before answering. It costs output tokens and latency.
| Raise effort for | Keep it low for |
|---|---|
| Planning a multi-file change | Mechanical renames and moves |
| Debugging where the cause is not obvious | Writing the fourth similar test |
| Choosing between two designs | Formatting, boilerplate, glue |
| Anything where being wrong is expensive to detect | Anything the type checker verifies instantly |
The pattern worth internalising: high effort for deciding, low effort for doing. Plan with the dial up, in its own session. Execute with the dial down, against the plan.
Watch out
High effort does not fix a context problem. If the agent has the wrong files loaded, thinking harder about the wrong files produces a more elaborate wrong answer. Reach for the effort dial only after you are confident the context is right — which is the opposite of most people’s reflex.
Determinism where it counts
You cannot make the model deterministic, so put determinism in the environment instead. Pinned dependency versions, a check command with stable output, tests that do not depend on wall-clock time or network. The more deterministic your environment, the less the model's variance matters — because variance that breaks something now gets caught in seconds.
Try it
Take one moderately hard task and run it three times in fresh, identical sessions. Diff the three results. The spread you see is your honest picture of that task’s reliability — and it tells you whether you need a check, a tighter spec, or both.
Takeaways
- Same prompt, different output, is expected behaviour — not a bug and not a sign of a bad model.
- Reliability comes from checks and specs, never from a prompt that worked once.
- High effort for decisions, low effort for execution. Effort does not compensate for bad context.
When is re-rolling a failed attempt the right move, and when is it a trap?
Right when the task is cheap to run and cheap to verify — a small function with a test, a config change the type checker validates. A trap when verification is expensive: re-rolling a large refactor means reviewing a large diff repeatedly, and you will get progressively less careful each time. There, fix the spec or shrink the task instead.
A course by Pieter Zandbergen