# When can you trust an AI to grade another AI’s answers?

> In this field note, veridive explains when a language model can be trusted to grade another model’s answers. Narrow yes/no checks with the evidence in view work better than one-to-ten scores. It covers writing rubrics a judge can apply, calibrating against human reviewers, known biases toward length, position and its own style, and where people must stay.

Automated grading makes evaluation cheap enough to run on every change, but the judge is one more system to validate. Use narrow yes/no checks, calibrate against people, watch for known biases and keep people on the high-risk cases.

## Key takeaways

- A model can grade another model’s answers when the check is narrow, yes or no, and the evidence is in front of it.
- Specific checks, such as “does the cited passage contain the amount?”, beat one-to-ten quality scores.
- Calibrate the judge against two human reviewers on a sample, and again after any change to its model, prompt or rubric.
- Watch for known biases toward longer answers, the first option shown and its own style, and keep people on high-risk cases.

The evaluation set has a few hundred examples, and scoring the system’s answers by hand takes two reviewers the better part of a day. Prompt changes queue up waiting for review, and before long the team starts shipping on intuition. That is usually when someone suggests LLM-as-a-judge: a model that grades another model’s answers.

It works, within limits. Automated grading makes evaluation cheap enough to run on every change. But the judge is one more AI system, with its own errors and biases, so it gets the same treatment as the system it grades: a narrow job, tests against people, and people kept on the cases that matter.

## What is LLM-as-a-judge?

A language model used as a grader. It receives the question, the system’s answer, the source passages and sometimes a reference answer, applies a rubric, and returns a verdict with a short reason. Three forms are common:

- **Rubric grading:** one answer checked against written criteria.
- **Reference comparison:** the answer compared with the approved answer in the [evaluation set](https://veridive.com/insights/evaluation-sets-are-the-new-requirements/).
- **Pairwise comparison:** two answers to the same question, and the judge picks the better one, typically when comparing two prompts or models.

It replaces routine human grading. It can’t replace the decision about what good means, which stays with the people who own the work.

## Which checks can a model grade reliably?

Narrow ones, with the evidence in the context and a yes-or-no answer. A “7 out of 10” can’t be verified, compared across runs or acted on; a specific check can.

| Check | Suits a judge? |
|---|---|
| Does the answer cite the passage that contains the amount it states? | Yes: the evidence is in view |
| Does the answer include each required fact? | Yes, one fact per check |
| Is the reply in formal Turkish (siz)? | Yes, with an example of each form |
| Does the answer make a claim the passages don’t support? | Mostly; calibrate carefully |
| Is the amount exactly right? | No: compare it with code |
| Is this the right business decision? | No: that is a person’s call |
| Rate overall quality from one to ten | No: vague and prone to bias |

Exact values such as dates, amounts, categories and ids belong to code, not to a model: code is faster, cheaper and never has an off day.

## How do you write a rubric a judge can apply?

The same way you write one for people, with more discipline:

- **One check per question,** phrased so the answer is yes or no.
- **The evidence in the prompt:** the passages, the reference outcome and the rubric line, so the judge compares instead of remembering.
- **A pass and a fail example** for every borderline check.
- **A reason before the verdict,** quoting the evidence, so disagreements can be audited.
- **Structured output:** check id, verdict and quoted evidence.
- **Versions:** the judge’s model, prompt and rubric are versioned like code, and any change means recalibration.

The [evaluation set template](https://veridive.com/insights/llm-evaluation-template/) shows rubric lines written this way.

## How do you validate the judge against people?

Treat it as an experiment with a pass mark:

1. **Sample rows across tags,** hard ones included, and have two reviewers score them independently with the same rubric.
2. **Check the reviewers first.** If they often disagree, the rubric is the problem; fix it before involving the judge.
3. **Compare judge and reviewers per check,** not overall. A judge can be reliable on citations and weak on tone.
4. **Look at the direction of errors.** A judge that passes answers people would fail is worse than one that is too strict.
5. **Set a threshold per check.** Checks that meet it are automated; the rest stay with people.
6. **Repeat after any change** to the judge, and periodically on fresh production samples.

Take an illustrative example: a team grades its policy assistant with a judge that scores each answer from one to ten. A prompt change makes answers longer, and scores rise. The reviewers aren’t convinced: the new answers aren’t better, just longer, and some carry an extra unsupported sentence. A sample confirms it: where judge and reviewers disagree, the judge nearly always preferred the longer answer. The team replaces the score with five yes/no checks: the rule is stated, the exception is stated, each amount is cited to the passage that contains it, no claim lacks support, and the answer fits the requested length. Two reviewers score a sample independently, two checks are reworded where they disagreed, and the judge is recalibrated against their agreed scores. The length effect disappears, because length no longer earns points, and the padding now fails the unsupported-claim check.

## Which biases should you watch for?

Several are well known, and each has a countermeasure:

- **Length:** a preference for longer, more detailed answers. Counter it with yes/no checks and a check for concision.
- **Position:** in pairwise comparisons, a preference for the answer shown first, or second. Run both orders and count only consistent verdicts.
- **Self-preference:** a preference for its own style or its own model family. Use a judge from a different family, and reference-based checks.
- **Confident form:** fluent answers with citations look right. Ask specifically whether the cited passage contains the claim.
- **Language:** a judge can be less reliable in Turkish than in English, so calibrate per language.

> The judge is one more AI system. Test it like one.

## Where must people stay in the loop?

Wherever a grading error would be expensive:

- **High-risk cases,** such as refunds, HR or legal answers and commitments to customers, are graded by people.
- **A regular sample** of judge-graded cases goes to people every week, to catch the judge drifting.
- **The rubric and the reference answers** belong to people; the judge only applies them.
- **Go-live decisions** are signed off by a person, with the judge’s results as evidence.

## Automate one check first

Pick the check your reviewers spend the most time on, write it as a yes/no question with the evidence in view, and compare the judge with two people on a sample before trusting it. Setting up evaluation sets and thresholds is [data and AI foundations](https://veridive.com/services/data-ai-foundations/) work; keeping the checks running after launch falls to [AI reliability](https://veridive.com/services/ai-reliability/). For the wider test strategy, see [how to test an LLM application](https://veridive.com/insights/testing-llm-applications/).

## Frequently asked questions

### What is LLM-as-a-judge?

LLM-as-a-judge means using a language model to grade the output of another AI system against a rubric. The judge receives the question, the answer, the source passages and sometimes a reference answer, and returns a verdict with a short reason. It makes evaluation cheap enough to run on every change, but the judge itself must be validated against human reviewers.

### Can you trust an LLM to evaluate another LLM?

For narrow checks, yes, once it has been calibrated. A judge that answers specific yes/no questions with the evidence in view, such as whether a cited passage contains the stated amount, can agree closely with human reviewers, which you measure rather than assume. Broad quality scores are less reliable and prone to bias. Keep people on high-risk cases and a regular sample.

## Sources

1. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (Zheng et al.). arXiv. https://arxiv.org/abs/2306.05685
