Two kinds of hallucination
One is obvious and cheap. The other survives review and is the reason you keep getting burned.
Hallucination means confidently-wrong output. It is worth splitting into two flavours, because they have different causes, different tells, and different fixes.
Factuality: invented from training
The model states something about the world that is not true — a library that does not exist, a function signature from a version that never shipped, a config option someone proposed in an issue. This comes from parametric knowledge: what the model absorbed during training, frozen at its knowledge cutoff.
Post-cutoff libraries and APIs are the classic trap. A framework that had a major release after the cutoff will get confidently described in terms of its previous major version, and the code will look completely plausible.
The good news: this flavour usually fails loudly. The import errors, the type does not exist, the test blows up. Your check command catches it.
Put the truth in context instead of relying on the model's memory: - point the agent at the actual node_modules types / vendored source - paste the relevant section of the real docs - tell it the version, explicitly, in your project brief - make it verify against the installed package before writing
Faithfulness: drifting from context it has
The model states something that contradicts material already in its own window. It read your schema and then described a column that is not in it. It read your handler and then summarised behaviour the handler does not have.
This is the expensive flavour, for three reasons: it sounds grounded, so you trust it; it passes type checks, because it is usually a claim rather than code; and it gets worse as the session grows, precisely when you are most tired and least likely to check.
| Factuality | Faithfulness | |
|---|---|---|
| Source | Training data | Attention degradation over loaded context |
| Typical form | Code that does not compile | A confident summary that is subtly wrong |
| Caught by | Types, tests, imports | Only by a human checking against the source |
| Gets worse with | Newer libraries | Longer sessions |
| Fix | Load the real thing into context | Shorter sessions; verify claims against primaries |
The habit to build
Treat every agent statement about your code as a claim with a citation attached. If the agent says "the retry logic lives in queue/worker.ts and caps at five attempts", the useful response is not "great" — it is to open the file, or to ask the agent to quote the lines. Quoting is cheap and it collapses the failure mode: a model asked to quote either quotes correctly or visibly cannot find it.
Before you summarise, quote the exact lines you are basing each claim on, with file path and line numbers. If you cannot find a line that supports a claim, say so rather than inferring it.
Try it
Go back to your baseline answer from Lesson 1.5. Classify each wrong claim as factuality or faithfulness. Most people find the faithfulness count higher than they expected — and those are the ones that would have shipped.
Takeaways
- Factuality errors fail loudly and your check command catches them.
- Faithfulness errors pass review because they sound sourced; they grow with session length.
- Ask for quotes with file paths. A claim that cannot be quoted is a claim that was inferred.
Why does "use the latest version of the library" not fix factuality problems?
Because the model has no way to know what the latest version is — its knowledge stops at the cutoff, and the instruction just makes it more confident about the last version it saw. The fix is to put the actual installed types, source, or docs into context so the answer comes from contextual knowledge rather than parametric memory.
A course by Pieter Zandbergen