Lesson 8.1Lesson 8.1 · Judgment, Ethics & Control
Hallucination & Verification
An agent can be fluent, confident and wrong all at once, because it is built to produce plausible language rather than true statements - so a practical verification discipline, applied in proportion to what a mistake would cost, is the non-negotiable price of trusting any output that matters
The most dangerous agent output is not the one that looks wrong - it is the one that looks perfectly right and isn't.
Ask an agent for the fire-egress width required by the local code, and it may answer instantly, in clean professional language, citing a clause number and a figure - all of it fluent, confident, and completely fabricated. This is a hallucination: a plausible, well-formed statement that is simply not true. It is not a rare glitch or a sign of a broken tool; it is a predictable property of how these systems work. A language model generates the most likely next words given everything before them - it is optimised to sound right, not to be right - and when it does not know something, it does not fall silent or flag its uncertainty the way an honest junior would. It fills the gap with something that fits the pattern. The result reads exactly like a correct answer, because fluency and correctness are produced by the same machinery and the model cannot reliably tell them apart. Neither can you, by reading.
That is why verification is the ethical core of agentic practice, and the first lesson of this module. Everything else - the architect of record, IP and data, bias and liability - rests on a single discipline: you do not rely on agent output you have not verified, in proportion to what a mistake would cost. For a mood-board caption, a glance is enough. For a code clause, a structural number, a cost, a specification, a claim to a client, or anything that becomes a binding commitment, you check it against an authoritative source before it leaves your hands. This lesson is about making that discipline concrete: understanding why agents are confidently wrong, knowing what to check and when, and building verification into how you work so it is a habit, not an afterthought you remember only after the error has shipped under your name.
Fluent is not true. Triage by consequence. Check against the real source, never a second AI. Then rely.
Why agents are confidently wrong
To verify well you have to understand what you are verifying against, and the honest answer is unsettling: an agent's fluency tells you nothing about its accuracy. A large language model, the engine inside most agents, is trained to predict plausible continuations of text. Given a question, it produces the words that most convincingly fit - drawing on patterns in its training data, the tools it can call, and the context you gave it. When the true answer is well represented and the agent can look it up, the plausible continuation and the correct one usually coincide, and you get a right answer. But when the true answer is obscure, absent, or contradicted by the pattern, the model still produces a plausible continuation - and now the plausible answer and the correct one part company. The model does not know it has crossed that line. It has no separate faculty that says "here I am guessing." It produces the confabulation in exactly the same confident register as a fact.
This is why the failure mode is so treacherous for professionals. A hallucination is not garbled or hedged; it is clean, specific and authoritative. An agent will invent a clause number for a code that has none, quote a standard that does not say what it claims, produce a dimension that is internally consistent but wrong, or cite a reference that does not exist - all in the measured voice of an expert. The tone carries no signal. You cannot separate truth from confabulation by how sure the answer sounds, because the answer always sounds sure. Worse, agentic systems compound the risk: an agent that has hallucinated a fact early in a multi-step task will treat it as established and build on it, so a small early fabrication can propagate silently through research, calculations and documents.
None of this makes agents useless - it makes them powerful instruments that require an external check. The mistake is to treat a fluent answer as a verified one. The discipline is to treat every agent output as a draft claim whose truth is still open until you have confirmed it against something outside the agent - the actual code text, the real drawing, the manufacturer's datasheet, a second independent method, or your own trained judgement. That external check is the only thing that reliably separates a correct answer from a confident wrong one, and building the habit of applying it is the whole of this lesson.
Fluency is not truth. The model sounds exactly as sure when it is wrong as when it is right.
A verification discipline: proportion to consequence
Verifying everything to the same depth is impossible and unnecessary; the practical discipline is to verify in proportion to what a mistake would cost. Before you rely on any agent output, ask one triage question: if this were wrong and I acted on it, what happens? The answer sorts the work into tiers, and each tier gets a matching level of check.
At the top are the must-verify outputs - anything touching life safety, structural adequacy, code and regulatory compliance, cost commitments, contractual claims, or a statement made to a client or authority as fact. These are verified against an authoritative source every time, without exception, before they leave your hands. A quoted code clause is checked against the actual code text, not another agent. A structural or egress number is re-derived or confirmed by a qualified professional. A specified product's performance is confirmed against the manufacturer's current datasheet. A cost is checked against real rates. Here the standard is not "it looks right" but "I have confirmed it right," because the consequence of a confident error is a dangerous building, a legal exposure, or a broken promise - with your name on it.
In the middle are outputs where an error is costly but recoverable - a research summary that will shape a decision, a draft schedule, a first-pass analysis. These deserve a real check: verify the load-bearing claims and any specific figures, sample the rest, and treat the whole as a draft to critique rather than an answer to accept. At the bottom are low-stakes outputs - internal notes, brainstorms, first-draft prose you will rewrite - where a glance suffices and heavy verification would be waste. The skill is calibrating the tier correctly, and erring upward when unsure. Two habits make this reliable: ask for sources and check them (an agent's citation is itself a claim to verify, not proof), and verify against something outside the agent - never ask the same or another agent to confirm the first, because a second fluent voice agreeing is not evidence, only correlated confidence. Build the triage question into your workflow and it becomes automatic: goal in, work out, check sized to the stakes, then rely.
What to check, and how to check it
Beyond triage, verification has a practical anatomy - specific things that go wrong in agent output and specific ways to catch them. Learn the failure modes and the checks become fast.
Fabricated specifics. Agents invent citations, clause numbers, standards, product codes, statistics and precedents that sound exact because they are exact-looking - a plausible IS number, a real-sounding case study, a precise figure. Check: never accept a specific reference on the agent's word. Open the actual source. If a clause is cited, read the clause. If a study is quoted, find the study. A citation that cannot be located almost certainly does not exist. Numerical and unit errors. Agents make arithmetic slips, drop or confuse units (mm versus m, kN versus kg), and produce internally consistent but wrong calculations. Check: re-derive critical numbers by an independent method or tool, sanity-check magnitudes against known benchmarks, and confirm units explicitly. Stale or wrong-jurisdiction facts. An agent may confidently give a superseded code edition, an old rate, or a rule from the wrong country or state - a live risk in India, where the model may default to foreign norms. Check: confirm the edition, date and jurisdiction against the current governing document.
Misread inputs. When an agent works over your drawings, schedules or documents, it can misread a dimension, transpose a value, or miss a note. Check: spot-check its outputs against the source document, especially where it summarises or extracts. Silent scope drift. Over a long task an agent may quietly answer a slightly different question than you asked, or carry an early assumption into every later step. Check: read the output against your original goal and watch for propagated assumptions. The unifying method is simple and it is the one non-negotiable: check against an authoritative source outside the agent. The code text, the real drawing, the manufacturer's data, a second tool, a qualified colleague, your own trained knowledge - these are what confirm truth. An agent confirming its own or another agent's claim is not verification; it is two confident voices, which is exactly the trap. Make "where did this come from and can I confirm it independently?" the reflex you apply to anything that matters, and the discipline holds.
Check the citation exists. Re-derive the number. Confirm the edition and jurisdiction. Compare to the real drawing.
Building verification into how you work
A verification discipline that lives only in good intentions fails under deadline pressure - the moment you are busy is exactly when the unverified output slips through. So the goal is to make verification structural: built into your process, your prompts and your studio's habits, so that relying on unchecked output feels wrong rather than convenient.
Start with how you direct the agent. Ask for its sources and its reasoning, not just its answer - an output you can trace is one you can check, and an agent that must show where a claim came from is easier to catch when the source is invented. Ask it to flag its own uncertainty and to distinguish what it looked up from what it inferred; it will not be perfectly honest, but it surfaces some risk. Constrain it to grounded sources where you can - an agent working from a document you supplied, or retrieving from a trusted corpus, hallucinates less than one answering from memory (Module 2.3). None of this replaces your check; it makes your check faster and better targeted.
Then build the check into the workflow. Define, for each recurring task, what tier it sits in and what the verification step is - so "draft the specification" always carries "confirm each product against its current datasheet," and "check the layout against the code" always ends with "read the actual clauses." Keep the human verification step as a named, non-skippable stage, not an optional afterthought - this is human-in-the-loop by design (Module 6.3). In a studio, agree the tiers as a team so verification is a shared standard, not one person's caution, and make clear that the person who relies on an output owns having verified it. Finally, cultivate the mindset that makes all of it work: treat every agent output as a capable draft from a fast, fluent, unreliable assistant - genuinely useful, never authoritative. You are not being paranoid; you are being professional. The designer who verifies as a reflex captures the full productivity of agents while carrying none of the hidden liability, because nothing confident and wrong ever leaves their hands unchecked. That reflex is the foundation the rest of this module - responsibility, data, ethics - is built on.
Cited codes, clauses & standards
Any regulatory or code reference an agent quotes
Read the actual clause in the governing document; confirm the edition and the jurisdiction. A citation is a claim to verify, not proof. Never accept an IS/NBC reference on the agent's word.
Safety, structural & egress numbers
Any figure where an error affects safety
Re-derive by an independent method or have a qualified professional confirm. Check units explicitly (mm/m, kN/kg). Confident arithmetic is still sometimes wrong.
Specifications, products & costs
Anything that becomes a binding commitment
Confirm against the manufacturer's current datasheet and real rates before it reaches a client or contract. A fabricated product code reads exactly like a real one.
Independence of the check
The source you verify against
Verify against something OUTSIDE the agent - the real document, a second tool, a qualified human. One AI confirming another is correlated confidence, not evidence.
Workshop — run a hallucination hunt
The fastest way to internalise verification is to catch an agent being confidently wrong yourself. In this workshop you will deliberately probe an agent on facts you can check, tier a real task, and design the verification step you will actually use.
An AI agent/assistant, one authoritative source you can check against (a code, a datasheet), and a notebook. The point is the checking, not the asking.
Goal: see hallucination first-hand and build a verification habit you will keep Inputs: an AI agent or assistant + one real code/spec question you know or can look up + a notebook Time: ~45 minutes
- 1Ask an agent five specific factual questions from your practice where you can check the answer - e.g. a named IS/NBC clause and its content, a product's fire rating, a required dimension, a case-study fact, a statistic. Ask for the source each time.
- 2Verify every answer against the authoritative source (the actual code text, the real datasheet, the genuine study). Record each as CONFIRMED, WRONG, or FABRICATED (source does not exist), and note how confident the wrong ones sounded.
- 3Take one real recurring task from your work (e.g. draft a spec, summarise precedents, check a layout against rules) and assign it a verification TIER: must-verify, real-check, or glance - and justify the tier by what a mistake would cost.
- 4For that task, write the exact verification STEP you will always perform before relying on the output - what you check, against which source, and who is responsible for doing it.
- 5Write a short reflection: how many of the five answers were wrong or fabricated, whether tone gave any of them away, and one change you will make to how you work with agents from now on.
You’ll walk away with
A one-page log of your hallucination hunt (five answers, verified) plus a written verification step for one real task, tiered by consequence. Keep it as your studio's starting verification standard.
Three altitudes on the same idea
Read the band that fits you — or all three.
As the architect of record you carry the consequence of every hallucination that reaches a drawing, so verification is not diligence - it is the job. Set a hard rule: nothing touching life safety, structure, fire, egress, code compliance or a cost commitment is relied upon on an agent's word. Quoted clauses are read in the actual code; critical numbers are re-derived or confirmed by a qualified professional; specified products are checked against current datasheets; jurisdiction and edition are confirmed (agents love to import foreign norms). Make the verification step a named, non-skippable stage in your document workflow, and make clear across the office that whoever relies on an output owns having checked it. An unverified agent output is a liability with your seal on it.
Agents accelerate your research, specification and documentation - and will just as fluently invent a product code, a fire rating, a dimension or a supplier that does not exist. Anything that becomes a commitment - a specification, a performance claim, a cost, a statement to a client - gets checked against the real source before you rely on it: the manufacturer's current datasheet, the actual price, the genuine lead time. Treat every fluent answer as a draft to confirm, ask the agent for its sources and open them, and never let a second AI voice "confirm" the first. Lower-stakes work - mood ideas, first-draft copy - needs only a glance. Size the check to the stakes, and the confident wrong answer never leaves your studio.
The single most valuable habit you can build now is to distrust fluency - to feel the reflex "can I confirm this independently?" every time an agent hands you a confident answer. Agents will happily invent citations, standards, numbers and case studies in flawless professional prose, and the tone gives nothing away; only an external check separates truth from confabulation. Practise it deliberately: whenever you use an agent for a fact, open the actual source and confirm it - you will be startled how often the plausible citation is fiction. This protects your grades and your future licence, and it builds the deeper skill of understanding work well enough to judge whether it was done right. Wield agents to do more; verify like the professional you are becoming.
“Hallucinations are basically a solved problem now - the newest models are so accurate, and can search the web and cite sources, that you can trust their output without checking it yourself.”
Do it yourself
Reason it through - and where you can, test it against a real agent.
- 1In one sentence, why does an agent sound exactly as confident when it is wrong as when it is right?
- 2What is the single triage question that decides how hard you verify an output?
- 3Give one example each of a fabricated specific, a numerical error, and a wrong-jurisdiction fact an agent might produce.
- 4Why is asking a second AI to confirm the first NOT verification?
- 5For a task in your own work, name the tier it belongs to and the exact source you would check against.
The one line to carry out
Peer-reviewed journals & authoritative standards
- 01Hallucination (artificial intelligence) — Wikipedia — Hallucination (artificial intelligence), 2026.
- 02Large language model — Wikipedia — Large language model, 2026.
- 03Human-in-the-loop — Wikipedia — Human-in-the-loop, 2026.
- 04Retrieval-augmented generation — Wikipedia — Retrieval-augmented generation, 2026.
Verification is what you owe the work - but who owes it? The next lesson names the person who cannot hand their responsibility to any tool: the architect of record.
The author
Amogh N P
Architect, interior designer, and creative polymath. Studio Matrx began in his notebooks — his vision of design made honest, useful, and open to everyone. Its Academy is written and taught in his memory, and free, forever.
More about Amogh →