Proof of concept, pilot or MVP: which AI step are you taking?
A proof of concept asks whether a model can do the task, a pilot asks whether it works on live work, and production promises it will keep working. Confusing the three is how impressive demos die.
veridive6 min read
The demo went well. A model summarized twenty contracts, the legal team liked the summaries, and before the meeting ended someone asked when it could go live. That question skips at least one step.
An AI proof of concept, a pilot and an MVP are not one thing at three sizes. Each answers a different question, runs on different data and is signed off by someone different. Confusing them is how impressive demos die: the project skips a question, and the answer arrives later, in front of users.
What question does each step answer?
The MVP is missing from the table on purpose: it describes a scope, not a step.
| How the steps differ | Proof of concept | Pilot | Production |
|---|---|---|---|
| Question | Can a model do this task? | Does it work on live cases, for real users? | Will it keep working, affordably, with an owner? |
| Data | A sample of past cases | Live cases from approved sources | Every case in scope |
| Users | Builders and one or two experts | The workflow owner and a small group of users | Everyone in scope |
| Duration | Days | Weeks | As long as the system runs |
| Exit criteria | A bar agreed in advance, met on a representative sample | Written acceptance criteria met on live work | None: it is monitored until it is retired |
| Signed off by | The sponsor funding the next step | The workflow owner | The workflow owner, at a go-live review |
What does a proof of concept prove, and what doesn’t it?
A good proof of concept proves three things quickly and cheaply: the model can read your kind of input, it can produce the output you need, and the cost per case is plausible.
It proves nothing about the cases your team actually receives, the systems the work runs through or the bill at full volume. The weak point is the sample. Proofs of concept usually run on examples someone chose: clean, typical, easy to find. Nobody picks their worst contract for a demo. That is the demo trap, and it isn’t dishonesty. It is selection.
Whoever picks the examples picks the result.
Three habits make a proof of concept worth its name:
- Sample, don’t choose. Pull cases at random from a normal period, then add the awkward ones on purpose.
- Set the bar first. Agree before the run which result justifies a pilot and which stops the project.
- List what it didn’t test. Document types, languages and integrations left out become the pilot’s first questions.
What makes a pilot a real pilot?
Here is an illustrative case. A legal team’s proof of concept summarized twenty contracts picked by a senior lawyer. The summaries were accurate, and a pilot was approved. Within days of live work, three problems surfaced:
- Scans. Older contracts arrived as scanned PDFs with stamps and handwritten changes, and text extraction garbled the clauses that mattered most.
- Annexes. Payment terms sat in annexes filed as separate documents, so the summaries missed the commercial terms.
- Amendments. Later amendments had changed prices and notice periods, and the summaries reported terms that no longer applied.
The proof of concept wasn’t wrong; it answered a narrower question than the one everyone heard. Surfacing the rest is the pilot’s job. So a real pilot runs on live cases as they arrive, with the workflow owner defining good and signing off. Its acceptance criteria are written before it starts: quality by case type, time and cost per case, and the errors that are never acceptable. Real users work with the output in their routine, and data protection is involved from the first week. A “pilot” on a fixed test set in a sandbox is just a longer proof of concept.
A careful pilot often begins in shadow mode, where the system works on live cases but nobody acts on its output yet.
Where does an MVP fit?
An MVP, a minimum viable product, is the smallest version real users can use for real work. In AI delivery it isn’t a separate step; it is the scope of the pilot: one team, one case type, one channel, with just enough integration and a review screen.
The risk is in what gets cut. Under deadline pressure, “minimum” starts to mean no evaluation set, no approval step and no logging. Those are the parts that show where an AI system fails, and its failures hide in cases you haven’t seen yet. Cut scope, never safeguards.
What are the exit criteria for each step?
Each step ends in a written decision:
- Proof of concept: the bar is met on a representative sample, the cost per case fits the business case, and the untested questions are listed. Outcomes: pilot, redesign or stop. An early no is cheap and useful.
- Pilot: acceptance criteria met on live work, users comfortable with the review step, data and security reviews closed. Outcomes: go live, go live with limits (one team or case type first), extend for a stated reason, or stop. The pilot-to-production checklist lists the eight conditions we check with the owner.
- Production: no exit, only upkeep: monitoring, regression checks, a cost ceiling and a runbook, until someone decides to retire the system.
Which step can you safely skip?
The proof of concept, sometimes. If the task is well understood, such as extracting fields from a familiar form, and you already have a strong evaluation set of real cases with expert-approved answers, the question is no longer whether a model can do it. It is whether it works here, which is a pilot question.
Keep the proof of concept when the input is unusual (poor scans, handwriting, mixed Turkish and English, heavy jargon), when nobody has tried the task on your data, or when a pilot needs expensive integrations before anyone knows the approach works.
Never skip the pilot for work that touches customers, money or commitments.
Name the step you are on
Write one sentence about your current AI project: the question it answers right now. If that sentence is “can a model do this?”, it is a proof of concept, whatever the slides call it.
If you are unsure, that is what a Discovery Sprint settles: a baseline, a data and access review, risks and a pilot plan with acceptance criteria. The ways in compare it with Pilot to Production, and how we work shows the steps in between.
Ask an assistant about this note