# How to design a review screen people can decide from in one look.

> In this field note, veridive explains how to design an AI review screen that lets a person decide in one look. It covers the anatomy of the screen, evidence shown next to each claim, uncertainty as reasons and bands instead of raw scores, one-click actions, structured override reasons, and testing the screen with the people who use it.

The review screen decides whether AI saves time. Show the recommendation, the evidence behind it, what is uncertain and what will happen on approval, so a person can decide without redoing the work.

## Key takeaways

- A review screen should let a person decide from one look: summary, recommendation, cited evidence, flags and the consequence of each action.
- Show uncertainty as reasons and bands calibrated on the evaluation set, not as raw probability scores.
- Make routine approval one click, keep friction for consequential actions, and capture edits and override reasons as structured data.
- Design and test the screen with the people who use it, on real cases, counting how often they leave it to check something.

Watch someone review an AI recommendation on a poorly designed screen. They read the suggestion, open the order system in another tab, then the policy document, then the customer’s history, and rebuild the case from scratch before clicking approve. The system did the work, and the screen made them do it again.

The review screen decides whether AI saves time. A good one shows the recommendation, the evidence behind it, what is uncertain and what will happen on approval, so a person can decide in one look without redoing the work.

## What does a reviewer need to see?

Six things, from top to bottom:

1. **A case summary.** Who, what and which order or document, in two or three lines, with the original request one click away.
2. **The recommendation.** One proposed decision or draft, in the team’s own terms.
3. **Evidence with citations.** Each fact the recommendation rests on, linked to its source: the order line, the policy paragraph, the photo.
4. **Flags.** Anything unusual or uncertain: missing data, conflicting sources, a repeat claim, an amount near a limit.
5. **Consequences.** What approve actually does (issues a refund, sends an email, posts an entry), and where edit and escalate lead.
6. **Three actions.** Approve, edit or escalate, always in the same place.

What stays off the screen matters as much: the full list of retrieved passages, internal scores and the model’s step-by-step reasoning. They belong in the audit log, one click away for the curious, not where the decision is made.

The test of done: a reviewer can explain the decision to a colleague using only the screen. The failure to fear most is a hidden consequence: someone approves what they think is a draft, and it goes straight to the customer.

## How should evidence be shown?

Next to the claim it supports, not as a list of documents at the bottom. Quote the exact passage and point to the paragraph, row or clause, with the version of the document, so checking takes seconds. A citation should open the source at that place, and only for people allowed to see it.

Show structured facts as fields from the source system (order date, amount, return window), filled in by code rather than paraphrased by the model. A number read from the system is as reliable as the system; a number retyped by a model needs checking.

Show negative evidence too. “No previous returns on this account” and “no policy exception applies” tell the reviewer what was checked. Without them, a careful reviewer checks again.

## How do you show uncertainty without misleading numbers?

A raw score such as “confidence 0.87” looks precise and says little. Scores a model reports about itself are often poorly calibrated, and people read them as a promise of accuracy that nobody tested.

Use bands and reasons instead:

- **Bands defined by evidence.** For example routine, check and decide, where each band is set by measured results on the evaluation set for that case type, not by the model’s own score.
- **Specific reasons.** “The photo doesn’t show the damage described” or “two policy paragraphs apply and they conflict” tells the reviewer where to look.
- **What would change the answer.** “If the customer confirms the item is unused, this becomes a full refund” turns a vague doubt into one question to check.

Re-check the bands after launch. A “routine” band that people override often is not routine.

> Show why the system is unsure, not how unsure it claims to be.

## Which actions should take one click?

Approve, when the recommendation is right: one click or one key, and the next case loads. Edit in place, with the draft already filled in, rather than in a separate form. Escalate with one click and a short required reason, into a named queue.

Consequential actions are the exception. Anything above a threshold or impossible to undo gets a confirmation that repeats the consequence (the amount, the recipient) before it runs. Speed is the point for routine cases; friction is the point for expensive ones. Be careful with bulk approval: if you offer it, limit it to one case type in the routine band, and sample those cases afterwards.

Watch the other risk as well: when approval is effortless and the system is usually right, approval can turn into a reflex. How to keep review meaningful is the subject of [why reviewers stop checking AI output](https://veridive.com/insights/automation-bias-in-ai-review/).

## How do you capture why a person disagreed?

As data, at the moment of disagreement, without slowing the reviewer down:

- **Edits.** Store every change as the difference between the draft and what was finally used.
- **Override reasons.** On every override, one click on a short fixed list (wrong source, missing data, policy changed, tone, other), plus optional free text.
- **Context.** Keep the case type, the evidence shown and the prompt and model versions with each record.

Don’t ask for reasons on approvals, and never make the reason a form. Who reads these records, and how they turn into fixes, is covered in [building a feedback loop from overrides](https://veridive.com/insights/ai-feedback-loop/).

## How do you test a review screen with its users?

Design it with the people who will use it, on real cases rather than demo cases, including the hard ones. Three measures tell you most of what you need: time to decision, whether the decision was right (known-answer cases from the evaluation set help here), and how often reviewers leave the screen to check something. Every “let me just check…” in a think-aloud session is a missing field. Test where the work happens, too: a screen that works in a meeting room can fail at a store counter or a busy contact-center desk.

In the illustrative [returns engagement](https://veridive.com/work/returns-intelligence/) on our work page, the review screen was designed with the returns team during weekly demos. Each case arrives with order and product data, a photo check, the policy paragraph behind the recommendation and a confidence indication, and the team confirms, edits or escalates in one click. Refunds above a threshold always need a person, and weekly reports compare recommendations with final decisions, which keeps the screen under test after launch.

## Watch three reviewers

Watch three reviewers use your current screen on real cases and count the tabs they open. Each one is evidence the screen should have shown. Review screens are part of every system we build as [custom AI software](https://veridive.com/services/custom-ai-software/), and the review experience has its own owner, a product designer, in the pod described in [how we work](https://veridive.com/approach/).

## Frequently asked questions

### What should an AI review screen show?

It should show a short case summary, the system’s recommendation, the evidence behind it with citations to the exact source, flags for anything unusual or uncertain, and what will happen on each action. The reviewer should be able to approve, edit or escalate from the same screen, without opening other systems to rebuild the case.

### How should AI confidence be shown to users?

As reasons and bands rather than raw probabilities. A score such as 0.87 invites false precision and is often poorly calibrated. A band such as routine, check or decide, defined by measured results on the evaluation set, plus a specific reason like “the photo doesn’t show the damage described”, tells the reviewer where to look.

### How do you capture feedback from human reviewers of AI output?

Record every edit as the difference between the draft and what was finally used, and ask for a reason on every override from a short fixed list, such as wrong source, missing data, policy changed or tone, with optional free text. Make it one click, never a form, and store it with the case type so patterns can be found.
