RAG or fine-tuning: which one does your use case need?
Retrieval gives a model knowledge it can cite and you can update daily; fine-tuning changes how it behaves. Most business knowledge problems need retrieval first, and fine-tuning only when evaluation shows a gap prompts can’t close.
veridive6 min read
“Can we train the model on our documents?” comes up early in many AI conversations. It usually means something reasonable: we want answers that reflect our policies, our products and our way of working. Training, or fine-tuning, is rarely the way to get that.
The two techniques change different things. Retrieval-augmented generation (RAG) changes what a model can see when it answers: passages from your approved documents, which it can cite and which you can update daily. Fine-tuning changes how a model behaves: its format, its style or its skill at a narrow task. Most business knowledge problems need retrieval first, and fine-tuning only when an evaluation shows a gap that better prompts can’t close.
What does each approach actually change?
Fine-tuning continues a model’s training on your own examples of inputs and desired outputs, so its weights (the internal parameters that shape its responses) shift toward that behavior. Retrieval leaves the model untouched and puts the relevant passages in front of it at the moment it answers; RAG, explained for business teams walks through how.
| Aspect | Retrieval (RAG) | Fine-tuning |
|---|---|---|
| What changes | What the model sees when it answers | How the model behaves |
| Data needed | Current, owned documents and their access rights | Many reviewed examples of the exact input and output you want |
| Update speed | As soon as a changed document is re-indexed | Only after a new training run, evaluation and release |
| Citations | Each claim can point to its passage | None: a trained-in fact has no source to show |
| Permissions | Filtered per user at search time | What was trained in is available to every user |
| Cost profile | A pipeline and index to run; longer prompts per call | Data preparation and training, repeated for each change; can lower cost per call |
| Main risks | Missed or outdated passages | Stale knowledge, confident errors, overfitting to the examples, dependence on one base model |
When is retrieval the right answer?
Whenever the problem is knowledge: the model needs facts it wasn’t trained on, those facts change, people need to check them, or different people may see different things. Policies, procedures, contracts, product data and technical documentation all fall here.
Take an illustrative case: an HR policy assistant for a group with several legal entities. Policies change a few times a year, differ by entity, and some documents, such as salary bands, are restricted. Retrieval handles all three: a new travel policy is live once it’s re-indexed, each employee sees only their entity’s version, and every answer cites its paragraph. A fine-tuned model would need re-training at every policy change, couldn’t show where an answer came from, and would know the salary bands for everyone.
That generalizes. Knowledge trained into a model can’t be cited, goes stale the day a document changes, and can’t respect document permissions. These are properties of the technique, not flaws a better training run fixes.
When does fine-tuning earn its cost?
When the problem is behavior, prompting hasn’t fixed it, and the evaluation set proves it. Typical cases:
- A fixed output format that must hold across many layouts and edge cases.
- A narrow, high-volume task where a smaller tuned model might match a larger model’s quality at lower cost and latency.
- Domain style or vocabulary that instructions and examples don’t keep consistent.
- Classification into your own categories when there are many and the boundaries are subtle.
A second illustrative case shows the pattern. A logistics team extracts the same fields from carrier delivery notes into a structure its ERP accepts. The fields never change; the layouts vary by carrier. With good instructions and examples, a large model gets most documents right, but the evaluation set shows persistent errors on two fields, and the cost per document is high at volume. That is a possible fine-tuning case: a smaller model trained on reviewed past extractions may hold the format more reliably at lower cost. “Possible” matters. The tuned model has to beat the prompted one on the same held-out examples before anyone switches.
What about better prompts and examples first?
Always first. The ladder has three steps, and you climb one only when the evaluation set shows a gap the current step can’t close:
- Prompts and examples. Clear instructions, the output format and a handful of worked examples. Cheapest to try, fastest to change.
- Retrieval. When failures come from missing, outdated or restricted knowledge.
- Fine-tuning. When the remaining failures are about behavior, or cost at volume justifies a smaller specialized model.
Each step should come with a sentence you can defend: which cases fail at the current step, why the next step would fix them, and how you will know. Skip the first step and you risk a training project for a problem a better prompt would have solved in an afternoon.
Fine-tuning changes how a model behaves. Retrieval changes what it knows when it answers.
What does each need in data, effort and upkeep?
- Prompting needs an evaluation set and prompts kept under version control. Upkeep: re-run the set when the model or the task changes.
- Retrieval needs a document pipeline, an index, permission sync and metadata. Upkeep is mostly content: owners retiring old versions and adding new ones.
- Fine-tuning needs training examples reviewed by people and kept apart from the evaluation set, a training setup or a provider’s fine-tuning service, and versioned model files. Upkeep: re-training when the task changes or the base model is retired, plus hosting if you run the tuned model yourself, a decision of its own covered in running your own model or using an API.
How do you decide with evidence?
- Build the evaluation set from real cases with approved answers.
- Run the prompted baseline and sort its failures into knowledge failures (missing or outdated facts) and behavior failures (format, style, consistency).
- Match the fix to the failure: retrieval for knowledge; prompts, then fine-tuning, for behavior.
- Compare on the held-out part of the set: quality, cost per case, speed and the upkeep you are signing up for.
- Write down the decision and what would reopen it, such as a new base model or a change in volume.
Knowledge or behavior?
Knowledge problem: retrieval. Behavior problem: prompts first, fine-tuning if the evaluation set proves you need it. Our data and AI foundations work builds the evaluation set and compares candidates on your own examples, and knowledge and document intelligence is where retrieval does most of its work.
Sources
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (Lewis et al.) arXiv arxiv.org/abs/2005.11401
Ask an assistant about this note