veridive is now an applied AI company. Looking for the answer engine?Looking for the answer engine? What happened

veridive TR Start a project Menu

Field notesEngineering

Evaluation sets are the new requirements document.

When the output is language, “done” is hard to define. A few hundred real examples with expert-approved answers make quality measurable — and turn model choice into an experiment instead of an opinion.

veridive6 min read

Traditional software projects begin with a requirements document. It lists what the system must do, and every item can be checked: the invoice screen shows the due date, the export includes the tax number, the report runs in under a minute. When the build is finished, someone ticks the boxes.

AI systems that produce language don’t fit that model. “The assistant answers policy questions accurately” is a sensible requirement that nobody can tick. Accurate for which questions? Judged by whom? What should happen when the policy is ambiguous, or the question is out of scope? Without answers, “done” becomes a matter of opinion, and projects drift between a demo that impressed and a launch nobody feels sure about.

The fix is an evaluation set: real examples with approved answers, and rules for judging the system against them. In our projects it does much of the work a requirements document used to do.

What is an evaluation set?

An evaluation set is a collection of real inputs from a workflow, each paired with the outcome a good expert would accept, plus a way to score the system’s output against it. For a returns workflow, one example might be a customer message, the order record and two photos, paired with the decision the returns team would make and the policy paragraph that justifies it. For a policy assistant, it might be an employee’s question paired with the correct answer and the document it comes from.

Written down, a single example has five parts:

  • Input: what the system receives, exactly as it arrives, attachments included.
  • Reference outcome: the answer or decision the owner would accept.
  • Evidence: the source that justifies it, such as a policy paragraph or an order field.
  • Tags: case type, difficulty and risk.
  • Scoring rule: how a reviewer, or an automated check, decides whether the system’s output matches.

A few hundred examples is usually enough to start. Coverage matters more than size: every common case type should be there, along with the rare cases that carry real risk.

How do you build one?

Start from live work, not from imagination. Pull recent cases from the systems where the work happens, across a normal spread of weeks, channels and teams. Then add the awkward ones on purpose: incomplete forms, contradictory information, questions the system should refuse or escalate, and cases where the right answer is “ask a person”.

The reference answers must come from the people who own the work, not from the builders. Where experts disagree, that disagreement is a finding: it usually means the policy is unclear, and a system cannot be more consistent than the rules it follows. Resolve it with the owner and write the decision down.

Tag each example with its case type, difficulty and risk. The tags show where the system is strong and where it fails, instead of a single average that hides both.

Finally, split the set. Keep a portion aside that is never used while tuning prompts or retrieval settings. That held-out part is the honest final test. The evaluation set template has the columns, tags and splits.

If you can’t say what a good answer looks like, no model can give you one.

How do you score language?

Some outputs can be scored exactly: a date, an amount, a category, a yes-or-no decision. Score those automatically. For free text, write a short rubric with the owner: which facts must appear, which statements must not, whether the source is cited correctly, whether the tone fits. A good rubric is one that two reviewers apply the same way.

Automated grading, including using a model to score answers against the rubric, makes it practical to run the set often. Treat the grader as one more system to validate: compare its scores with human reviewers on a sample, and keep people reviewing the high-risk cases.

How does it change the choice of model?

With an evaluation set, choosing a model becomes an experiment. Run each candidate against the same examples, for instance a large commercial model, a smaller and cheaper one, and an open-weight model that can run on your own servers, and compare quality, cost per case and speed. The question becomes practical: which is the least expensive option that meets the acceptance threshold on the cases that matter?

The same set tells you when to revisit the choice. When prices fall or a new model appears, re-run the set. The decision takes days instead of a debate.

How do you keep it alive?

An evaluation set is never finished. Every production failure that reaches a person becomes a new example. New products, policies and channels add new case types. We version the set, so results from different months can be compared, and review it with the owner every quarter.

Most importantly, it becomes the regression test, one layer of testing an LLM application. Every change to a model, a prompt or a data source runs the set before it reaches users. It is how you learn that an update has quietly made things worse before your customers do.

What goes wrong most often?

  • Only easy cases. A set built from typical examples scores well and tells you little.
  • Builders writing the answers. The team that builds the system should not decide what correct means.
  • Tuning on the test. If prompts are adjusted until the whole set passes, the score stops predicting live performance.
  • One number. An average hides the case types that fail, so report results by tag.
  • A set that never changes. Production brings new cases every month, and the set should follow.

What doesn’t it replace?

An evaluation set defines quality, not everything else. You still need ordinary requirements for integrations, permissions, response times, security and the screens people use. What changes is the center of gravity: the hardest question in an AI project, whether the output is good enough, now has a concrete, testable answer that the business owns.

Where to start

Pick one workflow. Pull a few dozen recent cases, ask the owner to write down the answers they would accept, and note where they hesitate. You will learn more about the workflow in that afternoon than in a month of requirement meetings, and you will have the start of the most useful document in your AI project.

Ask an assistant about this note

EngineeringEvaluationQuality

veridive

Field notes are written and reviewed by veridive. How we write them

Questions

Questions about this note

What is an evaluation set for AI?

An evaluation set is a collection of real examples from a workflow, each paired with the answer or decision an expert would accept, plus rules for scoring the system’s output. It works as the acceptance test for an AI system: it tells you whether quality is good enough to go live, and whether a change made things better or worse.

How many examples does an AI evaluation set need?

Usually a few hundred is enough to start, as long as they come from real work and cover the awkward cases, not only the typical ones. Coverage matters more than size: every important type of case and every consequential error should be represented. The set then grows as new edge cases appear in production.