veridive is now an applied AI company. Looking for the answer engine?Looking for the answer engine? What happened

veridive TR Start a project Menu

Field notesData & models

When can you trust an AI to grade another AI’s answers?

Automated grading makes evaluation cheap enough to run on every change, but the judge is one more system to validate. Use narrow yes/no checks, calibrate against people, watch for known biases and keep people on the high-risk cases.

veridive6 min read

The evaluation set has a few hundred examples, and scoring the system’s answers by hand takes two reviewers the better part of a day. Prompt changes queue up waiting for review, and before long the team starts shipping on intuition. That is usually when someone suggests LLM-as-a-judge: a model that grades another model’s answers.

It works, within limits. Automated grading makes evaluation cheap enough to run on every change. But the judge is one more AI system, with its own errors and biases, so it gets the same treatment as the system it grades: a narrow job, tests against people, and people kept on the cases that matter.

What is LLM-as-a-judge?

A language model used as a grader. It receives the question, the system’s answer, the source passages and sometimes a reference answer, applies a rubric, and returns a verdict with a short reason. Three forms are common:

  • Rubric grading: one answer checked against written criteria.
  • Reference comparison: the answer compared with the approved answer in the evaluation set.
  • Pairwise comparison: two answers to the same question, and the judge picks the better one, typically when comparing two prompts or models.

It replaces routine human grading. It can’t replace the decision about what good means, which stays with the people who own the work.

Which checks can a model grade reliably?

Narrow ones, with the evidence in the context and a yes-or-no answer. A “7 out of 10” can’t be verified, compared across runs or acted on; a specific check can.

CheckSuits a judge?
Does the answer cite the passage that contains the amount it states?Yes: the evidence is in view
Does the answer include each required fact?Yes, one fact per check
Is the reply in formal Turkish (siz)?Yes, with an example of each form
Does the answer make a claim the passages don’t support?Mostly; calibrate carefully
Is the amount exactly right?No: compare it with code
Is this the right business decision?No: that is a person’s call
Rate overall quality from one to tenNo: vague and prone to bias

Exact values such as dates, amounts, categories and ids belong to code, not to a model: code is faster, cheaper and never has an off day.

How do you write a rubric a judge can apply?

The same way you write one for people, with more discipline:

  • One check per question, phrased so the answer is yes or no.
  • The evidence in the prompt: the passages, the reference outcome and the rubric line, so the judge compares instead of remembering.
  • A pass and a fail example for every borderline check.
  • A reason before the verdict, quoting the evidence, so disagreements can be audited.
  • Structured output: check id, verdict and quoted evidence.
  • Versions: the judge’s model, prompt and rubric are versioned like code, and any change means recalibration.

The evaluation set template shows rubric lines written this way.

How do you validate the judge against people?

Treat it as an experiment with a pass mark:

  1. Sample rows across tags, hard ones included, and have two reviewers score them independently with the same rubric.
  2. Check the reviewers first. If they often disagree, the rubric is the problem; fix it before involving the judge.
  3. Compare judge and reviewers per check, not overall. A judge can be reliable on citations and weak on tone.
  4. Look at the direction of errors. A judge that passes answers people would fail is worse than one that is too strict.
  5. Set a threshold per check. Checks that meet it are automated; the rest stay with people.
  6. Repeat after any change to the judge, and periodically on fresh production samples.

Take an illustrative example: a team grades its policy assistant with a judge that scores each answer from one to ten. A prompt change makes answers longer, and scores rise. The reviewers aren’t convinced: the new answers aren’t better, just longer, and some carry an extra unsupported sentence. A sample confirms it: where judge and reviewers disagree, the judge nearly always preferred the longer answer. The team replaces the score with five yes/no checks: the rule is stated, the exception is stated, each amount is cited to the passage that contains it, no claim lacks support, and the answer fits the requested length. Two reviewers score a sample independently, two checks are reworded where they disagreed, and the judge is recalibrated against their agreed scores. The length effect disappears, because length no longer earns points, and the padding now fails the unsupported-claim check.

Which biases should you watch for?

Several are well known, and each has a countermeasure:

  • Length: a preference for longer, more detailed answers. Counter it with yes/no checks and a check for concision.
  • Position: in pairwise comparisons, a preference for the answer shown first, or second. Run both orders and count only consistent verdicts.
  • Self-preference: a preference for its own style or its own model family. Use a judge from a different family, and reference-based checks.
  • Confident form: fluent answers with citations look right. Ask specifically whether the cited passage contains the claim.
  • Language: a judge can be less reliable in Turkish than in English, so calibrate per language.

The judge is one more AI system. Test it like one.

Where must people stay in the loop?

Wherever a grading error would be expensive:

  • High-risk cases, such as refunds, HR or legal answers and commitments to customers, are graded by people.
  • A regular sample of judge-graded cases goes to people every week, to catch the judge drifting.
  • The rubric and the reference answers belong to people; the judge only applies them.
  • Go-live decisions are signed off by a person, with the judge’s results as evidence.

Automate one check first

Pick the check your reviewers spend the most time on, write it as a yes/no question with the evidence in view, and compare the judge with two people on a sample before trusting it. Setting up evaluation sets and thresholds is data and AI foundations work; keeping the checks running after launch falls to AI reliability. For the wider test strategy, see how to test an LLM application.

Sources

  1. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (Zheng et al.) arXiv arxiv.org/abs/2306.05685

Ask an assistant about this note

Data & modelsEvaluationQuality

veridive

Field notes are written and reviewed by veridive. How we write them

Questions

Questions about this note

What is LLM-as-a-judge?

LLM-as-a-judge means using a language model to grade the output of another AI system against a rubric. The judge receives the question, the answer, the source passages and sometimes a reference answer, and returns a verdict with a short reason. It makes evaluation cheap enough to run on every change, but the judge itself must be validated against human reviewers.

Can you trust an LLM to evaluate another LLM?

For narrow checks, yes, once it has been calibrated. A judge that answers specific yes/no questions with the evidence in view, such as whether a cited passage contains the stated amount, can agree closely with human reviewers, which you measure rather than assume. Broad quality scores are less reliable and prone to bias. Keep people on high-risk cases and a regular sample.