veridive is now an applied AI company. Looking for the answer engine?Looking for the answer engine? What happened

veridive TR Start a project Menu

Field notesData & models

An evaluation set template: columns, tags and a scoring rubric.

The working sheet behind an evaluation set: the columns, tag lists, rubric wording and splits to start from, described so a team can build it in an afternoon, with filled-in rows from a returns workflow and a policy assistant.

veridive7 min read

This is the working sheet behind Evaluation sets are the new requirements document. That note explains why real examples with approved answers replace much of a requirements document. This one skips the why and gives the template: the columns, the tag lists, the rubric wording and the splits, with a few filled-in rows.

A spreadsheet is enough to start. The structure matters more than the tool: stable ids, reference outcomes written by the people who own the work, tags you will actually report on, and scoring rules two reviewers apply the same way.

What columns does each example need?

One row per example, with twelve columns, plus the split column described further down.

ColumnWhat goes in itWatch for
idA stable identifier, such as RET-0107Never reuse or renumber
inputThe request exactly as it arrivedKeep typos and mixed languages
attachmentsPhotos, PDFs or records that came with itStore snapshots, not live links
reference outcomeThe answer or decision the owner would acceptWritten by the owner, not the builders
evidenceThe source behind it: document, version, passageEnough to find it again in seconds
case typeOne value from a fixed listOne per row
difficultyRoutine, tricky or edgeJudged by the owner
riskLow, medium or high: the cost of getting it wrongDefined in business terms
languageTurkish, English or mixedAdd script or dialect where relevant
scoring ruleExact match, a rubric id, or “must abstain”One rule per row
reviewerWho approved the reference outcomeA name, not a team
versionThe set version in which the row last changedBump it when the reference changes

Split the work by column. The builders pull inputs and attachments from the source systems and assign ids; the owner writes the reference outcome, the evidence and the risk; tags are agreed together. That division keeps the builders from deciding what correct means.

Which tags make results useful?

Tags exist to answer one question: where does the system fail? Keep each list short, mutually exclusive and agreed with the owner. A starting set:

  • Case type. For returns: damaged item, wrong item, duplicate order, changed mind, late delivery, warranty claim, suspected fraud. For a policy assistant: leave, travel, expenses, benefits, IT access.
  • Difficulty: routine, tricky, edge.
  • Risk: low, medium, high, defined in money, customer impact or compliance terms.
  • Language: Turkish, English, mixed.
  • Channel: email, chat, store, marketplace, call transcript.
  • Customer segment, such as consumer or business, new or returning. This is where uneven quality shows up.
  • Behavior flags: should abstain, should escalate, adversarial.

A tag earns its place only if a gap in it would change what you do next.

How do you write a rubric two reviewers apply the same way?

Write each rubric line as a yes/no check on something visible in the output. Four kinds of line cover most needs:

  • Must include: “States that visible transport damage needs no inspection before a replacement.” One fact per line.
  • Must not include: “Does not promise a refund before the item is received.” “Does not state any amount that is missing from the cited passage.”
  • Citation: “Cites the policy paragraph that contains the rule, and that paragraph supports every factual statement.”
  • Tone: “Uses formal address (siz) in Turkish.” “Does not blame the customer.”

For fields scored by exact match, write the normalization down as well: one date format, amounts rounded the same way, and whether “ödeme” and “odeme” count as the same answer.

Then calibrate. Two reviewers score the same thirty rows independently. Where they disagree, the line is ambiguous: rewrite it, or put a pass example and a fail example next to it. Repeat until disagreements are rare and explainable. Automated grading can apply the same lines at scale once it has been checked against these reviewers; LLM-as-a-judge covers how.

How do you split and version the set?

Add a split column with two values:

  • Development: rows you look at while tuning prompts, retrieval and models.
  • Held-out: rows used only for go/no-go decisions and major changes. Roughly a quarter of the set, drawn so every case type and language appears in both splits.

Two rules keep the held-out split honest: nobody tunes against it, and a held-out row that someone has studied moves to development and is replaced. Keep enough rows per case type that one failure doesn’t swing the result; a case type with only a handful of rows is a gap to fill, not a result to report.

Version the set like code. Adding rows is a minor version; changing a reference outcome or a rubric line is a major version, because old and new scores stop being comparable. Keep a changelog that says why each row was added: a production failure, a new policy, a new channel. When a policy changes, update the reference outcomes that depended on it and archive the old ones.

What do a few filled-in rows look like?

These rows are illustrative, invented for this note: two from a returns workflow and two from an HR policy assistant.

RowReference outcomeTagsScoring
RET-0107: “The sofa arrived with a torn seam. Photos attached. I want a replacement, not a refund.” Two photos, order record.Approve a replacement; visible transport damage needs no inspection; offer pickupDamaged item, medium risk, English, consumerOutcome match and rubric R-DMG
RET-0213: “Siparişi yanlışlıkla iki kez vermişim, birini iade edebilir miyim?” Order record.Approve return of the duplicate order with a free return label; reply in formal TurkishDuplicate order, low risk, Turkish, consumerOutcome match and tone line
POL-0031: “Can I carry unused leave over if I change teams mid-year?”Yes, up to the carry-over limit, with the new manager’s approval; cites the transfer section of the leave policyLeave, medium risk, EnglishRubric R-LEAVE and citation check
POL-0058: “What is the per-diem for a client visit abroad?” The per-diem annex is out of scope.Say no source was found and point to the travel deskTravel, should abstain, EnglishAbstention

The Turkish row asks, in effect, “I placed the order twice by mistake; can I return one?”

The template comes as a CSV starting point you can adapt: the twelve columns, the split column and a notes column, with these four rows filled in.

How should results be reported?

An average is where uneven quality goes to hide.

Report by tag, never only as one number:

  • A breakdown: pass rate by case type, risk and language, and by customer segment where it applies. A system that handles English email well and struggles with Turkish marketplace messages needs that shown, not averaged away.
  • Counts next to rates: a pass rate on four rows says far less than one on forty, and readers need to see which is which.
  • Separate measures: outcome match, rubric pass, citation support and correct abstention each get their own figure.
  • High-risk failures by name: every failed high-risk row listed by id, not folded into a rate.
  • Regressions: rows that passed in the previous run and fail now, even when the overall score rose.
  • Run details: the set version and the system version (model, prompt, retrieval settings), so any two runs can be compared.

That report is the evidence behind two of the go-live conditions on our checklist: acceptance criteria agreed in writing, and an evaluation set built from real examples.

Build it with the owner

Copy the twelve columns into a spreadsheet, add the split column, and fill in thirty real rows with the owner in one sitting. Building evaluation sets and acceptance thresholds is part of our data and AI foundations work, and the screen where people check the system’s proposals is covered in how to design a review screen.

Ask an assistant about this note

Data & modelsEvaluationTemplates

veridive

Field notes are written and reviewed by veridive. How we write them

Questions

Questions about this note

What columns should an LLM evaluation dataset have?

At minimum: a stable id, the input exactly as it arrived, any attachments, the reference outcome an expert approved, the evidence behind it, tags for case type, difficulty, risk and language, the scoring rule, the reviewer who approved the row, and the set version. Add a split column so held-out rows are never used for tuning.

How do you write an evaluation rubric for AI answers?

Write short yes/no checks that describe the output, not impressions: facts that must appear, statements that must not, whether the cited passage supports each claim, and tone rules such as formal address. Give a pass and a fail example for borderline lines, then have two reviewers score the same rows independently and rewrite any line they disagree on.