How to test an LLM application, layer by layer.
Test the deterministic parts like ordinary software, the language parts with evaluation sets, and the joins between them. A single end-to-end accuracy number hides where things actually break.
veridive6 min read
A team reports one number to its steering group: how often the assistant answers correctly on the evaluation set. After a change the number drops, and nobody can say why. It might be the retrieval index, a prompt edit, a date parser, a provider’s model update or a new batch of documents. The single number hides where things actually break.
Test an LLM application the way it is built: the deterministic parts like ordinary software, the language parts with evaluation sets, and the joins between them.
What is different about testing LLM applications?
Four things. The output is language, so there is rarely one exact answer to assert against. The output varies between runs, and the model behind it can change without a release on your side. The system is a pipeline of parsing, retrieval, prompt assembly, the model call, validation and integration, and an error in any stage looks like “the AI got it wrong”. And tests that call a model cost money and time.
Most of the application is still ordinary code, though. This plan keeps each kind of test where it is cheapest and most informative:
| Layer | What it tests | How | When |
|---|---|---|---|
| Unit | Parsing, prompt assembly, validators, business rules | Ordinary tests, model calls mocked | Every commit |
| Component | Retrieval and extraction on their own | Labeled queries and documents | Every change to that component |
| Evaluation set | Answer quality on real cases | Reference answers, rubric, repeated runs | Every prompt, model or source change, and nightly |
| End-to-end | The whole flow through real integrations | A few scripted cases in a test environment | Before each release |
| Adversarial | Injection, data leaks, approval bypass | Attack cases with expected refusals | With the evaluation set |
Which parts can be tested like ordinary code?
More than people expect:
- Input handling. Date and amount parsing, encoding, and case conversion for the Turkish İ and ı, which naive lowercasing gets wrong.
- Prompt assembly. The right instructions, examples and retrieved passages end up in the prompt, in order and within the length budget. Snapshot tests catch accidental changes.
- Validators and business rules. Schema checks, cross-field sums, thresholds.
- Plumbing. Tool wrappers, permissions, retries, timeouts and error handling.
In unit tests, mock the model: replace each call with recorded or canned responses, including malformed JSON, an empty reply and a timeout. The tests stay fast, free and repeatable, and they check what your code does with whatever the model returns.
How do you test retrieval on its own?
With a labeled set of queries and the passages that should come back (RAG, explained shows where retrieval sits), written with the content owner, including questions that have no answer in the sources. For each query, ask two things: did the right passage appear among the top results, and did anything come back that this user isn’t allowed to see?
Measuring retrieval separately tells you which half of the system to fix. Downstream, a retrieval miss and a generation error look identical.
If the right passage never came back, no prompt will fix the answer.
Re-run the retrieval tests on every change to chunking, embeddings, the index or the source documents.
Where do evaluation sets fit?
In the middle, testing the language parts on real cases with answers the business owner approved. How to build one is covered in evaluation sets are the new requirements document. For testing, three points matter. Score exact fields exactly and free text with a rubric. Report results by case type against the baseline, not as one average. And if a model grades the answers, validate the grader against people first, as described in when you can trust an AI to grade another AI.
How do you deal with output that varies between runs?
Reduce the randomness where the task allows, but don’t count on identical outputs; even at the lowest settings they aren’t promised. Then design the tests around variation:
- Repeated runs. Run each evaluation case several times. A case that passes only sometimes is a finding, usually an ambiguous instruction or a genuinely ambiguous case.
- Property checks. Assert what must be true rather than exact wording: the amount equals the order total, the answer cites a paragraph from the right policy, no other customer’s data appears.
- Tolerances from measured noise. Run the current version several times to see how much its score moves by chance. A change passes if it stays within that margin of the baseline on every important case type.
What runs on every change, and what runs nightly?
Put gates in the build pipeline. Unit tests with mocked models run on every commit. Any change to a prompt, retrieval setting, model or source runs the component tests and a fast slice of the evaluation set, and it can’t merge if an important case type falls below its threshold. Nightly, the full evaluation set runs with repeats, alongside the adversarial cases and the end-to-end script. Nightly runs also catch changes on the provider’s side, when nothing in your code moved.
Evaluations cost model calls: cases times runs times calls per case, plus grading. Keep the fast slice small, run the full set when it adds information, and budget for it like any other build infrastructure.
Take an illustrative policy assistant that answers employees’ travel and expense questions with citations (the numbers are invented):
- Unit: the date parser handles “31.12” and “end of quarter”; prompt assembly stays within budget; citations link to the paragraph. Model calls are mocked.
- Component: a hundred and twenty labeled questions check that the right paragraph comes back and that nothing from restricted HR files does.
- Evaluation set: three hundred real questions with the policy owner’s answers, scored for correctness and citation support, three runs each.
- Adversarial: requests for a colleague’s expense claims, and instructions to ignore the approval limit.
- Gates: unit and component tests on every commit, a fifty-case slice on every prompt or retrieval change, and every night the full set, the adversarial cases and a ten-question end-to-end script.
Split the single number
If your application has only an end-to-end accuracy number, split it: mock the model in your unit tests and build a retrieval test set. The next time a score drops, you will know which layer to open. An evaluation set and regression checks come with every custom AI software system we build, and AI reliability keeps them running after launch.
Sources
- OWASP Top 10 for Large Language Model Applications OWASP Gen AI Security Project genai.owasp.org/llm-top-10
- Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1 National Institute of Standards and Technology (NIST) nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf
Ask an assistant about this note