Why AI makes things up, and how to design so it gets caught.
You can’t remove hallucinations entirely, but you can make them rarer and visible: ground answers in approved sources, allow “I don’t know”, check names and numbers against systems of record, and put a person where a wrong answer is expensive.
veridive6 min read
You can’t train, prompt or buy your way to an AI system that never makes things up. You can make it happen less often, and make it visible when it does, so it is caught before it reaches a customer, a contract or a payroll run.
That is a design problem more than a model problem. Four moves do most of the work: ground answers in approved sources, let the system say “I don’t know”, check names and numbers against the systems that own them, and put a person wherever a wrong answer is expensive. Then measure what still slips through.
What is a hallucination, and why does it happen?
A hallucination is a statement an AI system presents as fact that its sources don’t support: a policy clause nobody wrote, a date that appears nowhere, a citation to a paragraph that says something else.
It happens because a language model generates the most plausible continuation of the text it was given. It has no built-in sense of which of its sentences rest on evidence, so when the answer isn’t in front of it, plausible text fills the gap. The risk rises in conditions you can spot in advance: the answer isn’t in the sources, the question assumes something false, the answer strings together many specific facts, or the instructions push the system to always answer.
Which kinds are most dangerous at work?
| Kind | What it looks like | Why it hurts |
|---|---|---|
| Invented fact | A deadline, fee or clause that exists in no document | It reads as authoritative, so nobody checks |
| Wrong or misattributed citation | A real document cited for a claim it doesn’t make | The receipt looks valid, so the claim borrows its credibility |
| One wrong detail | A correct answer with one invented number, name or date | Reviewers check the first lines, then skim |
| Confident extrapolation | A rule for one group applied to another, stated as fact | It sounds like reasoning, so it passes as judgment |
The third deserves the most design effort.
A plainly wrong answer gets caught. One wrong detail in a right answer gets trusted.
How does grounding reduce them?
Grounding means the system answers from approved passages retrieved for the question, not from what the model absorbed in training; RAG, explained for business teams walks through the mechanics. It lowers the rate in three ways:
- The facts are in front of the model. Most invention fills a gap, and a good passage removes the gap.
- The instructions narrow the job. Answer only from these passages, quote where it matters, and say so when they don’t cover the question.
- Each claim carries its source. When every statement links to a passage, an unsupported one stands out as a claim without a receipt.
Grounding can’t be better than the search behind it. If the right passage wasn’t retrieved, the model is back to filling gaps, which is why retrieval errors are the first place to look. Even with the right passage, a model can misread it, merge it with another or add a helpful-sounding detail. Grounding makes hallucinations rarer, not impossible.
Why should “I don’t know” be an allowed answer?
Many hallucinations are a system obeying an instruction to be helpful: if every question must get an answer, gaps get filled. Make abstaining a normal outcome in three places.
- In the instructions: when the passages don’t answer the question, say so, show the closest passages and name who to ask.
- In the interface: “no source found” looks like an answer, not an error.
- In the evaluation set: include questions the documents can’t answer, and score “I don’t know” as the correct response to them.
Watch the opposite failure too. A system that abstains on questions it could answer sends people back to email. Track both rates. Answers with receipts covers what an honest “no source found” should look like.
Which checks catch the rest?
Here is an illustrative case. An employee asks how much notice they must give before resigning. The assistant answers in three clear sentences and cites the HR policy. The process is right: written notice to the manager, a copy to HR, a handover plan. It also says the notice period is four weeks. The cited passage mentions no period at all; in this invented company, notice periods are set in each employment contract and vary by grade. Nobody wrote “four weeks” anywhere.
A claim check would have caught it. After the answer is generated, the system extracts its specific values (durations, amounts, dates, names) and looks for each one in the cited passages. “Four weeks” isn’t there, so the answer is held and regenerated without the figure, pointing to the contract and to HR instead.
That check belongs to a short list that catches most of what grounding misses:
- Values against sources. Numbers, names, dates and amounts in the answer must appear in the cited passage.
- Values against systems of record. An order total, a customer name or an invoice date comes from the ERP or CRM, inserted by the software, never retyped by the model.
- Citation checks. The cited passage exists, is the current version and contains the claim. An automated grader can run this at scale once it has been checked against people; see LLM-as-a-judge.
- A person where it’s expensive. Refunds and payments, anything sent to a customer or a regulator, decisions about an individual employee, contract terms and safety instructions get a review with the evidence side by side, as in our guardrails.
How do you measure the rate?
On the evaluation set, before launch and after every change:
- Split each answer into claims, one factual statement each.
- Label every claim as supported by the cited passage, supported only by another source, or unsupported.
- Count the unsupported-claim rate: the share of answers with at least one unsupported claim, with the claims listed for review.
- Report it by case type. Policy lookups, number-heavy questions and questions with no answer in the documents behave differently, and an average hides the difference.
- Weigh by severity. An unsupported amount matters more than an unsupported pleasantry.
In production, review a sample of live answers the same way, and add every reported hallucination to the evaluation set so the regression test catches it next time.
Mark every claim
Take twenty real questions, including five the documents can’t answer, and mark every claim in the answers as supported or not. That afternoon shows which of the four moves you need first. Grounded answers with a source for every claim are the core of knowledge and document intelligence, and the evaluation behind them is part of data and AI foundations.
Sources
- Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1 National Institute of Standards and Technology (NIST) nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf
- LLM09:2025 Misinformation OWASP Gen AI Security Project genai.owasp.org/llmrisk/llm092025-misinformation
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (Lewis et al.) arXiv arxiv.org/abs/2005.11401
Ask an assistant about this note