# Proof of concept, pilot or MVP: which AI step are you taking?

> In this field note, veridive explains how a proof of concept, a pilot and production differ: the first tests whether a model can do a task on samples, the second whether it works on live cases against written criteria, and production whether it keeps working. It covers exit criteria, where an MVP fits and when to skip steps.

A proof of concept asks whether a model can do the task, a pilot asks whether it works on live work, and production promises it will keep working. Confusing the three is how impressive demos die.

## Key takeaways

- A proof of concept tests whether a model can do the task; a pilot tests live work; production promises it keeps working.
- Hand-picked examples prove little about live work: sample cases at random, add the awkward ones and set the bar before testing.
- A real pilot runs on live cases with the workflow owner, real users and acceptance criteria written before it starts.
- Skip a proof of concept only for a well-understood task with a strong evaluation set; never skip the pilot.

The demo went well. A model summarized twenty contracts, the legal team liked the summaries, and before the meeting ended someone asked when it could go live. That question skips at least one step.

An AI proof of concept, a pilot and an MVP are not one thing at three sizes. Each answers a different question, runs on different data and is signed off by someone different. Confusing them is how impressive demos die: the project skips a question, and the answer arrives later, in front of users.

## What question does each step answer?

The MVP is missing from the table on purpose: it describes a scope, not a step.

| How the steps differ | Proof of concept | Pilot | Production |
|---|---|---|---|
| **Question** | Can a model do this task? | Does it work on live cases, for real users? | Will it keep working, affordably, with an owner? |
| **Data** | A sample of past cases | Live cases from approved sources | Every case in scope |
| **Users** | Builders and one or two experts | The workflow owner and a small group of users | Everyone in scope |
| **Duration** | Days | Weeks | As long as the system runs |
| **Exit criteria** | A bar agreed in advance, met on a representative sample | Written acceptance criteria met on live work | None: it is monitored until it is retired |
| **Signed off by** | The sponsor funding the next step | The workflow owner | The workflow owner, at a go-live review |

## What does a proof of concept prove, and what doesn’t it?

A good proof of concept proves three things quickly and cheaply: the model can read your kind of input, it can produce the output you need, and the cost per case is plausible.

It proves nothing about the cases your team actually receives, the systems the work runs through or the bill at full volume. The weak point is the sample. Proofs of concept usually run on examples someone chose: clean, typical, easy to find. Nobody picks their worst contract for a demo. That is the demo trap, and it isn’t dishonesty. It is selection.

> Whoever picks the examples picks the result.

Three habits make a proof of concept worth its name:

- **Sample, don’t choose.** Pull cases at random from a normal period, then add the awkward ones on purpose.
- **Set the bar first.** Agree before the run which result justifies a pilot and which stops the project.
- **List what it didn’t test.** Document types, languages and integrations left out become the pilot’s first questions.

## What makes a pilot a real pilot?

Here is an illustrative case. A legal team’s proof of concept summarized twenty contracts picked by a senior lawyer. The summaries were accurate, and a pilot was approved. Within days of live work, three problems surfaced:

- **Scans.** Older contracts arrived as scanned PDFs with stamps and handwritten changes, and text extraction garbled the clauses that mattered most.
- **Annexes.** Payment terms sat in annexes filed as separate documents, so the summaries missed the commercial terms.
- **Amendments.** Later amendments had changed prices and notice periods, and the summaries reported terms that no longer applied.

The proof of concept wasn’t wrong; it answered a narrower question than the one everyone heard. Surfacing the rest is the pilot’s job. So a real pilot runs on live cases as they arrive, with the workflow owner defining good and signing off. Its acceptance criteria are written before it starts: quality by case type, time and cost per case, and the errors that are never acceptable. Real users work with the output in their routine, and data protection is involved from the first week. A “pilot” on a fixed test set in a sandbox is just a longer proof of concept.

A careful pilot often begins in [shadow mode](https://veridive.com/insights/ai-shadow-mode-testing/), where the system works on live cases but nobody acts on its output yet.

## Where does an MVP fit?

An MVP, a minimum viable product, is the smallest version real users can use for real work. In AI delivery it isn’t a separate step; it is the scope of the pilot: one team, one case type, one channel, with just enough integration and a review screen.

The risk is in what gets cut. Under deadline pressure, “minimum” starts to mean no evaluation set, no approval step and no logging. Those are the parts that show where an AI system fails, and its failures hide in cases you haven’t seen yet. Cut scope, never safeguards.

## What are the exit criteria for each step?

Each step ends in a written decision:

- **Proof of concept:** the bar is met on a representative sample, the cost per case fits the business case, and the untested questions are listed. Outcomes: pilot, redesign or stop. An early no is cheap and useful.
- **Pilot:** acceptance criteria met on live work, users comfortable with the review step, data and security reviews closed. Outcomes: go live, go live with limits (one team or case type first), extend for a stated reason, or stop. The [pilot-to-production checklist](https://veridive.com/insights/ai-pilot-to-production-checklist/) lists the eight conditions we check with the owner.
- **Production:** no exit, only upkeep: monitoring, regression checks, a cost ceiling and a runbook, until someone decides to retire the system.

## Which step can you safely skip?

The proof of concept, sometimes. If the task is well understood, such as extracting fields from a familiar form, and you already have a strong evaluation set of real cases with expert-approved answers, the question is no longer whether a model can do it. It is whether it works here, which is a pilot question.

Keep the proof of concept when the input is unusual (poor scans, handwriting, mixed Turkish and English, heavy jargon), when nobody has tried the task on your data, or when a pilot needs expensive integrations before anyone knows the approach works.

Never skip the pilot for work that touches customers, money or commitments.

## Name the step you are on

Write one sentence about your current AI project: the question it answers right now. If that sentence is “can a model do this?”, it is a proof of concept, whatever the slides call it.

If you are unsure, that is what a Discovery Sprint settles: a baseline, a data and access review, risks and a pilot plan with acceptance criteria. The [ways in](https://veridive.com/services/#ways-in) compare it with [Pilot to Production](https://veridive.com/services/custom-ai-software/), and [how we work](https://veridive.com/approach/) shows the steps in between.

## Frequently asked questions

### What is the difference between an AI proof of concept and a pilot?

A proof of concept tests whether a model can do a task at all, usually on a sample of past cases with the builders watching. A pilot tests whether the system works on live cases as they arrive, used by the people who do the work, against acceptance criteria written in advance. A proof of concept ends in a decision to pilot; a pilot ends in a decision to go live.

### Can you skip the proof of concept and go straight to an AI pilot?

Yes, when the task is well understood and you already have a strong evaluation set: a few hundred real cases with expert-approved answers, including the awkward ones. The open question is then whether the system works in your workflow, which a pilot answers. Keep the proof of concept when inputs are unusual, the task is untested on your data or the integrations are expensive.

### Where does an MVP fit in an AI project?

An MVP, or minimum viable product, is the smallest version real users can use for real work. In an AI project it usually describes the scope of the pilot: one team, one case type and one channel, with just enough integration and a review screen. Keep the scope minimal but keep the evaluation set, the approval step and the logging, because they show where the system fails.
