An evaluation set template: columns, tags and a scoring rubric.
The working sheet behind an evaluation set: the columns, tag lists, rubric wording and splits to start from, described so a team can build it in an afternoon, with filled-in rows from a returns workflow and a policy assistant.
veridive7 min read
This is the working sheet behind Evaluation sets are the new requirements document. That note explains why real examples with approved answers replace much of a requirements document. This one skips the why and gives the template: the columns, the tag lists, the rubric wording and the splits, with a few filled-in rows.
A spreadsheet is enough to start. The structure matters more than the tool: stable ids, reference outcomes written by the people who own the work, tags you will actually report on, and scoring rules two reviewers apply the same way.
What columns does each example need?
One row per example, with twelve columns, plus the split column described further down.
| Column | What goes in it | Watch for |
|---|---|---|
| id | A stable identifier, such as RET-0107 | Never reuse or renumber |
| input | The request exactly as it arrived | Keep typos and mixed languages |
| attachments | Photos, PDFs or records that came with it | Store snapshots, not live links |
| reference outcome | The answer or decision the owner would accept | Written by the owner, not the builders |
| evidence | The source behind it: document, version, passage | Enough to find it again in seconds |
| case type | One value from a fixed list | One per row |
| difficulty | Routine, tricky or edge | Judged by the owner |
| risk | Low, medium or high: the cost of getting it wrong | Defined in business terms |
| language | Turkish, English or mixed | Add script or dialect where relevant |
| scoring rule | Exact match, a rubric id, or “must abstain” | One rule per row |
| reviewer | Who approved the reference outcome | A name, not a team |
| version | The set version in which the row last changed | Bump it when the reference changes |
Split the work by column. The builders pull inputs and attachments from the source systems and assign ids; the owner writes the reference outcome, the evidence and the risk; tags are agreed together. That division keeps the builders from deciding what correct means.
Which tags make results useful?
Tags exist to answer one question: where does the system fail? Keep each list short, mutually exclusive and agreed with the owner. A starting set:
- Case type. For returns: damaged item, wrong item, duplicate order, changed mind, late delivery, warranty claim, suspected fraud. For a policy assistant: leave, travel, expenses, benefits, IT access.
- Difficulty: routine, tricky, edge.
- Risk: low, medium, high, defined in money, customer impact or compliance terms.
- Language: Turkish, English, mixed.
- Channel: email, chat, store, marketplace, call transcript.
- Customer segment, such as consumer or business, new or returning. This is where uneven quality shows up.
- Behavior flags: should abstain, should escalate, adversarial.
A tag earns its place only if a gap in it would change what you do next.
How do you write a rubric two reviewers apply the same way?
Write each rubric line as a yes/no check on something visible in the output. Four kinds of line cover most needs:
- Must include: “States that visible transport damage needs no inspection before a replacement.” One fact per line.
- Must not include: “Does not promise a refund before the item is received.” “Does not state any amount that is missing from the cited passage.”
- Citation: “Cites the policy paragraph that contains the rule, and that paragraph supports every factual statement.”
- Tone: “Uses formal address (siz) in Turkish.” “Does not blame the customer.”
For fields scored by exact match, write the normalization down as well: one date format, amounts rounded the same way, and whether “ödeme” and “odeme” count as the same answer.
Then calibrate. Two reviewers score the same thirty rows independently. Where they disagree, the line is ambiguous: rewrite it, or put a pass example and a fail example next to it. Repeat until disagreements are rare and explainable. Automated grading can apply the same lines at scale once it has been checked against these reviewers; LLM-as-a-judge covers how.
How do you split and version the set?
Add a split column with two values:
- Development: rows you look at while tuning prompts, retrieval and models.
- Held-out: rows used only for go/no-go decisions and major changes. Roughly a quarter of the set, drawn so every case type and language appears in both splits.
Two rules keep the held-out split honest: nobody tunes against it, and a held-out row that someone has studied moves to development and is replaced. Keep enough rows per case type that one failure doesn’t swing the result; a case type with only a handful of rows is a gap to fill, not a result to report.
Version the set like code. Adding rows is a minor version; changing a reference outcome or a rubric line is a major version, because old and new scores stop being comparable. Keep a changelog that says why each row was added: a production failure, a new policy, a new channel. When a policy changes, update the reference outcomes that depended on it and archive the old ones.
What do a few filled-in rows look like?
These rows are illustrative, invented for this note: two from a returns workflow and two from an HR policy assistant.
| Row | Reference outcome | Tags | Scoring |
|---|---|---|---|
| RET-0107: “The sofa arrived with a torn seam. Photos attached. I want a replacement, not a refund.” Two photos, order record. | Approve a replacement; visible transport damage needs no inspection; offer pickup | Damaged item, medium risk, English, consumer | Outcome match and rubric R-DMG |
| RET-0213: “Siparişi yanlışlıkla iki kez vermişim, birini iade edebilir miyim?” Order record. | Approve return of the duplicate order with a free return label; reply in formal Turkish | Duplicate order, low risk, Turkish, consumer | Outcome match and tone line |
| POL-0031: “Can I carry unused leave over if I change teams mid-year?” | Yes, up to the carry-over limit, with the new manager’s approval; cites the transfer section of the leave policy | Leave, medium risk, English | Rubric R-LEAVE and citation check |
| POL-0058: “What is the per-diem for a client visit abroad?” The per-diem annex is out of scope. | Say no source was found and point to the travel desk | Travel, should abstain, English | Abstention |
The Turkish row asks, in effect, “I placed the order twice by mistake; can I return one?”
The template comes as a CSV starting point you can adapt: the twelve columns, the split column and a notes column, with these four rows filled in.
How should results be reported?
An average is where uneven quality goes to hide.
Report by tag, never only as one number:
- A breakdown: pass rate by case type, risk and language, and by customer segment where it applies. A system that handles English email well and struggles with Turkish marketplace messages needs that shown, not averaged away.
- Counts next to rates: a pass rate on four rows says far less than one on forty, and readers need to see which is which.
- Separate measures: outcome match, rubric pass, citation support and correct abstention each get their own figure.
- High-risk failures by name: every failed high-risk row listed by id, not folded into a rate.
- Regressions: rows that passed in the previous run and fail now, even when the overall score rose.
- Run details: the set version and the system version (model, prompt, retrieval settings), so any two runs can be compared.
That report is the evidence behind two of the go-live conditions on our checklist: acceptance criteria agreed in writing, and an evaluation set built from real examples.
Build it with the owner
Copy the twelve columns into a spreadsheet, add the split column, and fill in thirty real rows with the owner in one sitting. Building evaluation sets and acceptance thresholds is part of our data and AI foundations work, and the screen where people check the system’s proposals is covered in how to design a review screen.
Ask an assistant about this note