veridive is now an applied AI company. Looking for the answer engine?Looking for the answer engine? What happened

veridive TR Start a project Menu

Field notesData & models

RAG, explained for business teams: how AI answers from your documents.

Retrieval-augmented generation is less about the model than about the library: which documents are approved, who may see them, how they are found and how every answer points back to its source. Most quality problems start there.

veridive6 min read

Ask a general-purpose chatbot what your company’s travel policy allows and it will answer anyway, from what it absorbed in training about other companies. A well-built company assistant does something different: it searches your approved documents first and answers from the passages it found, with a link to the paragraph.

That pattern is called retrieval-augmented generation, or RAG. It is simpler than the name suggests, and less about the model than people expect. Most of the work, and most of the quality problems, sit in the library: which documents are approved, who may see them, how they are found and how each answer points back to its source.

What is RAG, in one paragraph?

Retrieval-augmented generation is a pattern in which an AI system first retrieves passages relevant to a question from approved documents, then asks a language model (the kind of AI that reads and writes text) to answer using only those passages, and shows which passage supports each statement. The model supplies the reading and the writing; your documents supply the facts. Nothing is trained into the model, so when a policy changes, you replace the document and the next answer follows it.

How does a question become a cited answer?

In an illustrative case, an employee asks the company assistant, “Can I book business class for a six-hour flight to a client meeting?”

  1. Question. The system knows who is asking from the company login, including their role and, where policies differ, their country or legal entity.
  2. Search. It searches only the documents this employee may open, for meaning (“long-haul flight”) and for exact terms (“business class”), and returns a shortlist: the flights section of the travel policy and an annex on trips paid for by clients.
  3. Passages. A handful of the best passages, not whole documents, go to the model, with instructions to answer only from them.
  4. Answer. In this invented policy, flights over five hours may be booked in business class with a director’s approval, so the answer says yes, with approval, and names the approval step.
  5. Citation. Each statement links to its passage, with the quoted sentence and the document’s version, so the employee can check it in seconds.

If nothing in the library answers the question, the right response is “I couldn’t find this in the documents you can access”, with the closest passages and the policy owner’s name. “No source found” is a valid answer, not a failure; answers with receipts covers how to design it.

What does RAG need from your documents?

An assistant can’t be more current or more consistent than its library. Before any model is involved, the library needs:

  • An approved scope: the document sets that are in, and those that stay out, such as drafts and personal folders.
  • An owner for each set, who says which version is current and settles it when two documents disagree.
  • One current version, because search finds old versions as easily as new ones.
  • Readable text: OCR (optical character recognition) for scans, and tables that keep their rows and columns.
  • Metadata: owner, effective date, status, audience and language, so search can filter and citations can show the version.
  • Permissions it can check, taken from the source system.

The last one is not negotiable. The assistant must never become a way around your access rights: people get answers only from documents they could open themselves, checked at search time, before the model sees any text. Access control for assistants has its own note.

Where do RAG systems usually go wrong?

Mostly before the model writes a word:

  • The wrong version was found. An expired policy still sits in a shared folder and gets retrieved with, or instead of, the current one.
  • The right passage was never found. The question uses different words, or contains a product code, a clause number or a name that meaning-based search handles badly. Turkish word forms make this harder.
  • The passage was cut badly. A rule lands in one passage and its exception in the next.
  • Structure was lost. A table flattened into a line of numbers, or a scan with no text layer.
  • The question needed clarifying. “What is the notice period?” depends on the country, the contract and the role, and the system guessed.
  • The model added something. Less often, the right passage was there and the answer still went beyond it.

Most wrong answers were decided at search time, before the model wrote a word.

So log the passages retrieved for every answer. If the right one wasn’t among them, changing the prompt or the model won’t help. Why document assistants give wrong answers walks through the diagnosis.

When is RAG the right approach, and when isn’t it?

RAG fits when knowledge lives in documents that change, when people need to check the source, and when different people may see different documents: policies, procedures, contracts, manuals. For neighboring problems, something else usually fits better.

If people need…A better fit is usually…
A figure from the ERP or the data warehouseA query or report on the system that owns the number
To find a form, not an answerBetter search: titles, metadata, synonyms
A fixed output format or a consistent styleBetter instructions and examples, then possibly fine-tuning
A judgment call, such as an exception to policyA person, with the relevant passages gathered for them

Two of these rows have their own notes: RAG or fine-tuning and assistant or better search. And if the documents don’t exist, or nobody owns them, the first project is a content project, not an AI one.

What should a business owner ask before building one?

Six questions. A team that can answer them is ready to build; a team that can’t has found its first piece of work.

  1. Sources. Which document sets are in scope, and which are explicitly out?
  2. Owners. Who owns each set, and who decides when two documents disagree?
  3. Freshness. How quickly must a changed document reach the answers, and how are old versions retired?
  4. Permissions. Is access checked at search time with the source system’s rules, and how fast do changes take effect?
  5. Citations. Does every answer point to the exact passage and version, and what does the user see when no source is found?
  6. Evaluation. Which real questions, with answers the owner has approved, will show that it works before launch and after every change?

The last one matters most over time. It turns “the assistant seems good” into a measurement, and it catches the day a document update or a model change makes things worse without anyone noticing.

One library, one owner

Pick one document set with frequent questions and a clear owner, such as HR policies, collect a few dozen real questions about it, and answer the six questions for that set alone.

Our data and AI foundations work covers the layers underneath: approved sources, permission-aware retrieval and evaluation. Knowledge and document intelligence shows where cited answers get used, and a short brief is the easiest way to talk through a specific library.

Sources

  1. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (Lewis et al.) arXiv arxiv.org/abs/2005.11401

Ask an assistant about this note

Data & modelsRAGCitations

veridive

Field notes are written and reviewed by veridive. How we write them

Questions

Questions about this note

What is RAG in simple terms?

RAG, short for retrieval-augmented generation, is a way for an AI system to answer from your own documents. It first searches an approved library for passages relevant to the question, then has a language model write an answer using only those passages, and shows which passage supports each statement. Nothing is trained into the model, so updating a document updates the answers.

Does a RAG assistant respect document permissions?

Only if it is built and tested to. A well-built RAG assistant filters its search by the access rights in the source system, so people get answers only from documents they could open themselves. The filter has to run before the model sees any text, stay in sync when access changes, and cover caches and logs as well as answers.

What do you need before building a RAG assistant?

You need an approved set of documents with a named owner, current versions clearly marked, text the system can read, and access rights it can check. You also need a few dozen to a few hundred real questions with approved answers and sources, so you can measure quality before launch and after every change to documents, prompts or models.