# Shadow mode: how to test AI on live work without risking it.

> In this field note, veridive explains shadow mode: running an AI system on live cases while people work as usual, then comparing its proposals with their decisions. It covers what must be in place first, how long to run it, measuring agreement by case type, reading disagreements with the owner and moving to assisted work.

Run the system on real cases while people work as usual, compare its proposals with what they decided, and let the agreement rate by case type decide what moves forward. It is the cheapest honest test there is.

## Key takeaways

- Shadow mode runs the system on live cases while people work as usual; nobody acts on its proposals.
- Keep people blind to the output, log every proposal with its versions, and check data protection before live cases flow.
- Report agreement by case type, and review disagreements with the owner: some are human slips or unclear policy.
- Move to assisted work one case type at a time, against exit criteria written before the shadow run started.

For a few weeks, the system reads every return request the team receives and writes down what it would decide. Nobody sees its answers, and the team works exactly as before. Every day, someone puts the two sets of decisions side by side.

That is shadow mode: the system runs on live work in parallel with the current process, and nothing it proposes reaches a customer or a record. It is the cheapest honest test there is. It uses the real case mix instead of a sample, it carries no operational risk, and it produces the number to decide on before anyone relies on the system: how often it agrees with the people who do the work, by type of case.

## What is shadow mode, and what is it for?

It answers a question a proof of concept can’t: does the system hold up on everything that actually arrives, including the cases nobody thought to put in the evaluation set? It also tests the plumbing at real volume: integrations, response times and cost per case. In a [real pilot](https://veridive.com/insights/ai-proof-of-concept-vs-pilot/), it usually comes first.

The discipline that makes it work is simple. In pure shadow mode, people don’t see the output. Once a suggestion is on screen, decisions start to lean toward it, and agreement stops being an independent measure. For the same reason, shadow mode can’t tell you how people will work with the output or whether the review screen is any good. Assisted work tests that later.

## What needs to be in place before it starts?

- **A pass on past cases.** The system has met the bar on the evaluation set. Shadow mode is not the place to discover that it can’t read the documents.
- **Read-only access.** It reads live inputs and writes nowhere: no records, no emails, no customers.
- **Logging.** One record per case: the input, the proposal with its evidence and confidence, the model, prompt and source versions, the person’s decision, timestamps and the case type. Without the versions, you can’t tell which change moved the numbers.
- **A data-protection check.** Live cases usually contain personal data, now processed by a new system and perhaps in a new place. Before the first case flows, ask your data protection officer about the legal basis, the processing location, how long logs are kept and who may read them.
- **Case types and matching rules.** Agree with the owner how cases are grouped and what counts as agreement: the same category, say, or an amount within a set tolerance.
- **Exit criteria, in writing.** Which agreement on which case types moves forward, and what stops the test.

Done looks like one complete record per case, collected automatically. The classic mistake is logging the system’s answer without the person’s decision, which turns every comparison into a manual project.

## How long should it run?

Set the length by cases, not by the calendar. Run until every case type in scope has enough comparisons to judge, and until the run has covered the periods that change the mix: a month-end, a campaign, a seasonal peak. For high-volume work that is often a few weeks.

Two mistakes pull in opposite directions. One is stopping when the overall number looks good, before the rare, expensive case types have appeared. The other is waiting indefinitely for a case type that arrives twice a month. If something is too rare to judge, it stays with people, and the rest moves on.

## What do you measure?

Agreement by case type comes first. Next to it: the cases the system declined or escalated, and whether that was right; errors by severity, since a wrong refund amount and a wrong tag are different failures; and response time and cost per case at real volume.

Here is an illustrative example from a returns desk. The numbers are invented to show how to read the table; they are not results.

| Case type | Agreed | Most disagreements | Next step |
|---|---|---|---|
| Unopened, inside the return window | 176 of 180 | Staff slips on dates | Move to assisted work |
| Damaged on arrival, photos attached | 78 of 90 | System accepted blurry photos | Add a photo-quality check; stay in shadow |
| Missing parts | 17 of 25 | Policy silent on partial refunds | Owner clarifies the policy; re-run |
| High-value or repeat returns | 11 of 12 | Too few cases to judge | Stays with people |

The overall figure, 282 of 307, looks healthy and hides every decision that matters. Each row gets its own answer.

## How do you read disagreements?

With the owner, one by one at first, before anyone blames the system. Disagreements come in four kinds:

1. **The system was wrong.** The case joins the evaluation set, and the cause gets fixed: a source, a rule or a prompt.
2. **The person was wrong.** A missed date, a forgotten clause. That is a finding about the current process, whose error rate was never zero.
3. **The policy is unclear.** Both answers are defensible. The owner decides and writes the rule down; a system can’t be more consistent than its rules.
4. **The person knew something the system didn’t.** A phone call, a note in another system. Connect that source, or keep the case type with people.

> A disagreement is a question, not a verdict.

Watch one trap: tuning prompts until agreement rises on the very cases you are measuring. Keep a slice of shadow cases that nobody tunes against, or the number stops predicting anything.

## How do you move from shadow to assisted work?

One case type at a time, against the exit criteria. In the table above, only the first row moves. In assisted work, people see the proposal and its evidence on a review screen and decide, so the measures change: override rate, time per case, and signs that approval is becoming a reflex. Keep a small blind sample, cases decided without the proposal on screen, so an independent comparison survives.

The illustrative [returns engagement](https://veridive.com/work/returns-intelligence/) on our work page follows this pattern: during the pilot, the assistant ran alongside the existing process on live cases, and decisions were compared daily. The results then feed the go-live review, alongside the [other conditions we check](https://veridive.com/approach/#checklist). After launch, a rising override rate in one case type is the earliest warning you will get, as the note on [monitoring metrics](https://veridive.com/insights/llm-monitoring-metrics/) explains. Letting a task act without approval is a later, separate decision, made [one narrow task at a time](https://veridive.com/insights/when-ai-agents-should-act-alone/).

## Where the comparison is cheap

Pick a workflow whose decisions are already recorded in a system, so the comparison costs almost nothing to collect. If you want a pilot designed this way from the first week, the Pilot to Production engagement in our [custom AI software](https://veridive.com/services/custom-ai-software/) work runs the pilot next to the current process on live cases, as [how we work](https://veridive.com/approach/) describes.

## Frequently asked questions

### What is shadow mode in AI testing?

Shadow mode is a test in which an AI system processes live cases in parallel with the existing process, while people keep working as usual and nobody acts on the system’s output. Its proposals are logged and compared with the decisions people actually made, broken down by case type, to show where the system is ready to assist and where it is not.

### How long should an AI shadow mode test run?

Until every case type in scope has enough comparisons to judge, and the run has covered the periods that change the case mix, such as a month-end or a campaign. For high-volume work that is often a few weeks. Case types too rare to judge in reasonable time stay with people rather than holding up the rest.

### What does agreement rate mean in shadow mode?

It is the share of cases in which the system’s proposal matches the decision a person made independently, under a matching rule agreed in advance, such as the same category or an amount within a set tolerance. It is only meaningful when people have not seen the proposal, and it should be reported by case type rather than as one average.
