Your baseline
Grade the answer you saved on the course home page. It tells you what your agent does when you give it no process at all.
On the course home page you asked your agent to explain your project, based only on files it had read. Open that answer now and mark it honestly.
Grade it against four questions
- Did it read, or did it guess? Scroll back through the session. Count the actual file reads. An answer that describes your architecture after reading a README and two filenames is parametric knowledge — a plausible description of a project like yours, not of yours.
- Is every specific claim checkable? Take three concrete statements — a file path, a function name, a data flow — and verify them. Note each one that is wrong. Wrong claims that contradict files the agent actually read are faithfulness hallucinations, and they are the expensive kind.
- What did it miss? Name the thing about your project that a new engineer most needs to know. Is it in the answer? Usually it is not, because it lives in convention rather than in code.
- How confident did it sound relative to how right it was? This gap is the thing you are learning to correct for. It does not close on its own.
Record the number
Put a line in your scratch log: how many of the specific claims were wrong, out of how many you checked. It is a crude metric and that is fine. You will run the same exercise at the end of Module 05, after you have given the agent a project brief, and the difference is the clearest measurement of what steering buys you.
## Baseline — cold exploration, no brief Reads before answering: 3 Specific claims checked: 6 Wrong: 2 (said state lives in Redux; we moved to Zustand 8 months ago) Missed entirely: the two-database split. Would break any agent that touched writes.
Watch out
Do not read a bad baseline as "this agent is not good enough". A cold agent with no brief, exploring an unfamiliar codebase, getting a third of its specifics wrong is normal. The point of the measurement is that almost all of that error is addressable by process, and you are about to learn the process.
Try it
Grade your baseline and log it. Then, separately, write down the one thing the agent missed that would have caused the most damage. That sentence is the first line of the project brief you write in Module 05.
Takeaways
- An unverified explanation is a hypothesis; count the file reads behind it.
- Confidence and accuracy are uncorrelated in model output. Calibrate yourself, not the model.
- Baseline now so you can measure what steering buys you later.
What is the difference between the agent inventing a library that does not exist and the agent misstating something in a file it just read?
The first is a factuality hallucination — it comes from parametric knowledge and is usually caught immediately, because the import fails. The second is a faithfulness hallucination: it drifted from context it actually had. That one passes review, because it sounds like it came from the code, and it gets worse as the session grows.
A course by Pieter Zandbergen