# Why document assistants give wrong answers, and where to look first.

> In this field note, veridive explains why document assistants give wrong answers. Most errors start in the library and the search, not the model: expired duplicates, flattened tables, passages split from their exceptions, search misses and missing filters. It offers a diagnostic table, named failure patterns and an order of fixes, starting with logging retrieved passages.

When a document assistant is wrong, the model is usually the last suspect. Most errors come from the library and the search: outdated or duplicate documents, broken tables, passages cut mid-thought, missing permissions or a question that needed clarifying.

## Key takeaways

- When a document assistant is wrong, check the library and the search first; the model is usually the last suspect.
- Log the retrieved passages for every answer: without them you can’t tell a retrieval error from a generation error.
- Expired duplicates, flattened tables and passages that split a rule from its exception are common causes, and prompting fixes none of them.
- Fix in order: sources, then parsing and chunking, then search, then prompts, and only then consider the model.

A user reports that the document assistant got the meal allowance wrong. The first instinct is to reword the prompt, or to try a larger model. Both are usually the wrong place to start.

When a document assistant is wrong, the model is the last suspect worth checking. The answer was usually lost earlier: an expired version in the library, a table that came apart in conversion, a passage cut between a rule and its exception, a search that never found the right paragraph, or a question that needed a clarifying question back. One habit makes all of this diagnosable: log the passages the system retrieved for every answer. Without that log, teams tune prompts in the dark.

## How do you tell a retrieval error from a generation error?

Open the retrieval log for the wrong answer and ask one question: was the passage containing the right answer among those sent to the model? ([RAG, explained for business teams](https://veridive.com/insights/rag-explained-for-business-teams/) describes the moving parts.)

- **No:** a retrieval error. The cause is in the documents, the parsing or the search, and no prompt will fix it.
- **Yes, but the answer contradicts or adds to it:** a generation error.
- **The right passage doesn’t exist:** a content gap. The correct answer was “no source found”.

Then work through the table.

| Symptom | Likely cause | First check | Fix |
|---|---|---|---|
| Quotes an outdated rule | Expired or duplicate version in the index | Source and version of the cited passage | Retire the old copy; filter on status and date |
| Wrong figure from a table | Table flattened in conversion | The passage text as indexed | Table-aware parsing, then re-index |
| States the rule, misses the exception | A chunk boundary between them | The neighboring chunks | Chunk by section; keep clauses whole |
| “No source found” for a known answer | Scan without text, or a wording mismatch | Is the text searchable? Do keywords hit? | OCR with checks; hybrid search; synonyms |
| Right topic, wrong country or entity | No metadata filter, or an ambiguous question | The filters applied and the user’s context | Add filters; ask a clarifying question |
| Right passage, wrong answer | Generation error | The passage against the answer, claim by claim | Tighter instructions; claim checks |
| Cites a document the user can’t open | Missing or stale permission filter | Access lists in the index against the source | Filter at retrieval; sync permissions |

## What goes wrong with the documents themselves?

- **The expired twin.** Old and new versions are both indexed, and search ranks whichever matches the wording better. Warning signs: citations with old effective dates, or different answers to the same question on different days. Fix: one current version in scope, the rest archived outside it, and status as a search filter.
- **The orphan set.** Nobody owns the documents, so contradictions stay unresolved. Warning sign: two current documents that disagree and nobody who can say which is right. Fix: an owner per set, with contradictions sent to them as policy questions.
- **The missing page.** The answer lives in an inbox or in someone’s head. Warning sign: repeated “no source found” on the same topic. Fix: a content backlog for the owner. The [document clean-up checklist](https://veridive.com/insights/prepare-documents-for-ai/) covers the rest.

An illustrative example: an assistant that quotes an expired travel policy. An employee asks for the daily meal allowance abroad and gets an amount that was correct two revisions ago. The retrieval log shows the cited passage came from “Travel policy – final.pdf” in a finance team folder, not the HR policy library: a copy saved there long ago and indexed when the folder was added to scope. The current policy was retrieved too, but ranked second, because the old copy repeated the question’s wording more closely. The fix needed no prompt change. The finance copy left the scope and was replaced with a link, expired documents were filtered out by status, and the question joined the evaluation set with the correct answer, so the regression test would catch a repeat.

## What goes wrong in parsing and chunking?

Parsing turns files into text; chunking splits the text into passages for search.

- **The flattened table.** A rate table becomes a run of numbers, and a real figure lands next to the wrong heading. Warning sign: correct-looking amounts attached to the wrong category. Fix: parse tables row by row with headers repeated in each row, and compare a sample with the originals.
- **The silent scan.** A PDF that is only an image has no text to search. Warning sign: documents that are never cited. Fix: OCR checked by eye on a sample, including Turkish characters.
- **The orphaned exception.** A fixed-length split puts “business class is allowed on long flights” in one passage and “except on trips billed to a client” in the next. Warning sign: answers that state the rule and miss “unless” or “except” (in Turkish, “hariç” or “ancak”). Fix: chunk along the document’s own sections and clauses, keep the heading with each chunk, and fetch the neighboring chunk when a clause continues.
- **Boilerplate in the middle.** Headers, footers and page numbers split sentences in two. Fix: strip them before chunking.

## What goes wrong in search?

- **Exact-term blindness.** Meaning-based search misses codes, clause numbers and names; [hybrid search](https://veridive.com/insights/hybrid-search-rag/) adds the keyword side.
- **Word-form misses.** Turkish suffixes turn one word into many strings unless documents and queries are analyzed the same way.
- **No filters.** Without metadata filters, an expired document, another entity’s policy or the wrong language competes on equal terms.
- **The wrong number of passages.** Too few, and the answer sits just below the cut; too many, and near-misses distract the model. Measure recall at the cut-off on the evaluation set.
- **The unasked question.** When a question maps to several documents that disagree by country, contract or role, ask which one applies instead of guessing.

## What goes wrong when the answer is written?

Less often than the rest, but it happens:

- **The helpful addition:** a detail the passage doesn’t contain. Fix: answer only from the passages, and check that every number and name appears in the source; see [how to reduce AI hallucinations](https://veridive.com/insights/how-to-reduce-ai-hallucinations/).
- **The merge:** two passages for different cases blended into one answer. Fix: cite per claim and state the conditions.
- **The missing “I don’t know”:** the system answers when it should abstain. Fix: allow abstention and test it with unanswerable questions.
- **Citation drift:** a real passage cited for a claim it doesn’t support. Fix: an automated citation check, calibrated against people.

## In what order should you fix things?

> Fix the library before the prompt, and the prompt before the model.

1. **Log retrieval for every answer:** the question, the filters, the passages and their scores.
2. **Fix the sources:** versions, duplicates, owners, missing documents.
3. **Fix parsing and chunking:** tables, scans, clause boundaries.
4. **Fix search:** hybrid retrieval, filters, reranking, the number of passages.
5. **Then adjust instructions and prompts.**
6. **Then, if the evaluation set still shows a gap, test another model.**

After each fix, re-run the evaluation set and add the failing question to it, so the fix stays fixed.

## Diagnose recent mistakes

Take the twenty most recent wrong answers users reported, pull their retrieval logs and sort each into retrieval error, generation error or content gap. If there are no logs, that is the first fix. The same diagnosis applies to any [knowledge and document intelligence](https://veridive.com/solutions/knowledge-document-intelligence/) system, and tracing and repairing these layers is everyday [data and AI foundations](https://veridive.com/services/data-ai-foundations/) work for us.

## Frequently asked questions

### Why does my RAG system give wrong answers?

Usually because the right passage never reached the model, or the wrong one did. Common causes are expired or duplicate documents in the index, tables and scans that lost their structure, passages that split a rule from its exception, search that misses codes or word forms, and questions that needed clarifying. Log the retrieved passages for each answer to see which one applies.

### How do you tell a retrieval error from a generation error?

Look at the passages the system retrieved for that answer. If the passage containing the right answer is not among them, it is a retrieval error, and changing the prompt or the model won’t fix it. If the right passage was there and the answer still contradicts or goes beyond it, it is a generation error.
