Agent Engineering
Module 02 · Concepts/Lesson 2.6/3 min

Non-determinism and effort

Why the same prompt gives different answers, and what the reasoning dial actually buys you.

Run the same prompt twice against the same model and you can get different output. This is non-determinism, and it has two sources: sampling during generation, and variation in how the provider serves the request (batching, hardware, routing). Even at temperature zero, providers do not generally guarantee bit-identical results.

What this means for your process

  • A success is not a guarantee. A prompt that worked once will not always work. If you need reliability, the reliability has to come from the check, not the prompt.
  • A failure is not a verdict. Before concluding the agent cannot do something, retry once in a fresh session. Surprisingly often the second attempt is fine, and you have learned that the task sits near the edge of what it does reliably — useful information.
  • Re-rolling is a legitimate technique, but only for cheap, verifiable tasks. Re-rolling a 40-minute refactor you then have to review three times is not a technique, it is a lottery.

Effort

Effort is a dial exposed by most current harnesses (sometimes called reasoning effort, thinking budget, or extended thinking). Turned up, the model generates a larger internal reasoning pass before answering. It costs output tokens and latency.

Raise effort forKeep it low for
Planning a multi-file changeMechanical renames and moves
Debugging where the cause is not obviousWriting the fourth similar test
Choosing between two designsFormatting, boilerplate, glue
Anything where being wrong is expensive to detectAnything the type checker verifies instantly

The pattern worth internalising: high effort for deciding, low effort for doing. Plan with the dial up, in its own session. Execute with the dial down, against the plan.

Watch out

High effort does not fix a context problem. If the agent has the wrong files loaded, thinking harder about the wrong files produces a more elaborate wrong answer. Reach for the effort dial only after you are confident the context is right — which is the opposite of most people’s reflex.

Determinism where it counts

You cannot make the model deterministic, so put determinism in the environment instead. Pinned dependency versions, a check command with stable output, tests that do not depend on wall-clock time or network. The more deterministic your environment, the less the model's variance matters — because variance that breaks something now gets caught in seconds.

Try it

Take one moderately hard task and run it three times in fresh, identical sessions. Diff the three results. The spread you see is your honest picture of that task’s reliability — and it tells you whether you need a check, a tighter spec, or both.

Takeaways

  • Same prompt, different output, is expected behaviour — not a bug and not a sign of a bad model.
  • Reliability comes from checks and specs, never from a prompt that worked once.
  • High effort for decisions, low effort for execution. Effort does not compensate for bad context.
When is re-rolling a failed attempt the right move, and when is it a trap?

Right when the task is cheap to run and cheap to verify — a small function with a test, a config change the type checker validates. A trap when verification is expensive: re-rolling a large refactor means reviewing a large diff repeatedly, and you will get progressively less careful each time. There, fix the spec or shrink the task instead.

A course by Pieter Zandbergen