Agent Engineering
Module 02 · Concepts/Lesson 2.5/3 min

Two kinds of hallucination

One is obvious and cheap. The other survives review and is the reason you keep getting burned.

Hallucination means confidently-wrong output. It is worth splitting into two flavours, because they have different causes, different tells, and different fixes.

Factuality: invented from training

The model states something about the world that is not true — a library that does not exist, a function signature from a version that never shipped, a config option someone proposed in an issue. This comes from parametric knowledge: what the model absorbed during training, frozen at its knowledge cutoff.

Post-cutoff libraries and APIs are the classic trap. A framework that had a major release after the cutoff will get confidently described in terms of its previous major version, and the code will look completely plausible.

The good news: this flavour usually fails loudly. The import errors, the type does not exist, the test blows up. Your check command catches it.

the fix
Put the truth in context instead of relying on the model's memory:
  - point the agent at the actual node_modules types / vendored source
  - paste the relevant section of the real docs
  - tell it the version, explicitly, in your project brief
  - make it verify against the installed package before writing

Faithfulness: drifting from context it has

The model states something that contradicts material already in its own window. It read your schema and then described a column that is not in it. It read your handler and then summarised behaviour the handler does not have.

This is the expensive flavour, for three reasons: it sounds grounded, so you trust it; it passes type checks, because it is usually a claim rather than code; and it gets worse as the session grows, precisely when you are most tired and least likely to check.

FactualityFaithfulness
SourceTraining dataAttention degradation over loaded context
Typical formCode that does not compileA confident summary that is subtly wrong
Caught byTypes, tests, importsOnly by a human checking against the source
Gets worse withNewer librariesLonger sessions
FixLoad the real thing into contextShorter sessions; verify claims against primaries

The habit to build

Treat every agent statement about your code as a claim with a citation attached. If the agent says "the retry logic lives in queue/worker.ts and caps at five attempts", the useful response is not "great" — it is to open the file, or to ask the agent to quote the lines. Quoting is cheap and it collapses the failure mode: a model asked to quote either quotes correctly or visibly cannot find it.

prompt
Before you summarise, quote the exact lines you are basing each claim on,
with file path and line numbers. If you cannot find a line that supports a
claim, say so rather than inferring it.

Try it

Go back to your baseline answer from Lesson 1.5. Classify each wrong claim as factuality or faithfulness. Most people find the faithfulness count higher than they expected — and those are the ones that would have shipped.

Takeaways

  • Factuality errors fail loudly and your check command catches them.
  • Faithfulness errors pass review because they sound sourced; they grow with session length.
  • Ask for quotes with file paths. A claim that cannot be quoted is a claim that was inferred.
Why does "use the latest version of the library" not fix factuality problems?

Because the model has no way to know what the latest version is — its knowledge stops at the cutoff, and the instruction just makes it more confident about the last version it saw. The fix is to put the actual installed types, source, or docs into context so the answer comes from contextual knowledge rather than parametric memory.

A course by Pieter Zandbergen