veridive is now an applied AI company. Looking for the answer engine?Looking for the answer engine? What happened

veridive TR Start a project Menu

Field notesDelivery

Proof of concept, pilot or MVP: which AI step are you taking?

A proof of concept asks whether a model can do the task, a pilot asks whether it works on live work, and production promises it will keep working. Confusing the three is how impressive demos die.

veridive6 min read

The demo went well. A model summarized twenty contracts, the legal team liked the summaries, and before the meeting ended someone asked when it could go live. That question skips at least one step.

An AI proof of concept, a pilot and an MVP are not one thing at three sizes. Each answers a different question, runs on different data and is signed off by someone different. Confusing them is how impressive demos die: the project skips a question, and the answer arrives later, in front of users.

What question does each step answer?

The MVP is missing from the table on purpose: it describes a scope, not a step.

How the steps differProof of conceptPilotProduction
QuestionCan a model do this task?Does it work on live cases, for real users?Will it keep working, affordably, with an owner?
DataA sample of past casesLive cases from approved sourcesEvery case in scope
UsersBuilders and one or two expertsThe workflow owner and a small group of usersEveryone in scope
DurationDaysWeeksAs long as the system runs
Exit criteriaA bar agreed in advance, met on a representative sampleWritten acceptance criteria met on live workNone: it is monitored until it is retired
Signed off byThe sponsor funding the next stepThe workflow ownerThe workflow owner, at a go-live review

What does a proof of concept prove, and what doesn’t it?

A good proof of concept proves three things quickly and cheaply: the model can read your kind of input, it can produce the output you need, and the cost per case is plausible.

It proves nothing about the cases your team actually receives, the systems the work runs through or the bill at full volume. The weak point is the sample. Proofs of concept usually run on examples someone chose: clean, typical, easy to find. Nobody picks their worst contract for a demo. That is the demo trap, and it isn’t dishonesty. It is selection.

Whoever picks the examples picks the result.

Three habits make a proof of concept worth its name:

  • Sample, don’t choose. Pull cases at random from a normal period, then add the awkward ones on purpose.
  • Set the bar first. Agree before the run which result justifies a pilot and which stops the project.
  • List what it didn’t test. Document types, languages and integrations left out become the pilot’s first questions.

What makes a pilot a real pilot?

Here is an illustrative case. A legal team’s proof of concept summarized twenty contracts picked by a senior lawyer. The summaries were accurate, and a pilot was approved. Within days of live work, three problems surfaced:

  • Scans. Older contracts arrived as scanned PDFs with stamps and handwritten changes, and text extraction garbled the clauses that mattered most.
  • Annexes. Payment terms sat in annexes filed as separate documents, so the summaries missed the commercial terms.
  • Amendments. Later amendments had changed prices and notice periods, and the summaries reported terms that no longer applied.

The proof of concept wasn’t wrong; it answered a narrower question than the one everyone heard. Surfacing the rest is the pilot’s job. So a real pilot runs on live cases as they arrive, with the workflow owner defining good and signing off. Its acceptance criteria are written before it starts: quality by case type, time and cost per case, and the errors that are never acceptable. Real users work with the output in their routine, and data protection is involved from the first week. A “pilot” on a fixed test set in a sandbox is just a longer proof of concept.

A careful pilot often begins in shadow mode, where the system works on live cases but nobody acts on its output yet.

Where does an MVP fit?

An MVP, a minimum viable product, is the smallest version real users can use for real work. In AI delivery it isn’t a separate step; it is the scope of the pilot: one team, one case type, one channel, with just enough integration and a review screen.

The risk is in what gets cut. Under deadline pressure, “minimum” starts to mean no evaluation set, no approval step and no logging. Those are the parts that show where an AI system fails, and its failures hide in cases you haven’t seen yet. Cut scope, never safeguards.

What are the exit criteria for each step?

Each step ends in a written decision:

  • Proof of concept: the bar is met on a representative sample, the cost per case fits the business case, and the untested questions are listed. Outcomes: pilot, redesign or stop. An early no is cheap and useful.
  • Pilot: acceptance criteria met on live work, users comfortable with the review step, data and security reviews closed. Outcomes: go live, go live with limits (one team or case type first), extend for a stated reason, or stop. The pilot-to-production checklist lists the eight conditions we check with the owner.
  • Production: no exit, only upkeep: monitoring, regression checks, a cost ceiling and a runbook, until someone decides to retire the system.

Which step can you safely skip?

The proof of concept, sometimes. If the task is well understood, such as extracting fields from a familiar form, and you already have a strong evaluation set of real cases with expert-approved answers, the question is no longer whether a model can do it. It is whether it works here, which is a pilot question.

Keep the proof of concept when the input is unusual (poor scans, handwriting, mixed Turkish and English, heavy jargon), when nobody has tried the task on your data, or when a pilot needs expensive integrations before anyone knows the approach works.

Never skip the pilot for work that touches customers, money or commitments.

Name the step you are on

Write one sentence about your current AI project: the question it answers right now. If that sentence is “can a model do this?”, it is a proof of concept, whatever the slides call it.

If you are unsure, that is what a Discovery Sprint settles: a baseline, a data and access review, risks and a pilot plan with acceptance criteria. The ways in compare it with Pilot to Production, and how we work shows the steps in between.

Ask an assistant about this note

DeliveryPilotsEvaluation

veridive

Field notes are written and reviewed by veridive. How we write them

Questions

Questions about this note

What is the difference between an AI proof of concept and a pilot?

A proof of concept tests whether a model can do a task at all, usually on a sample of past cases with the builders watching. A pilot tests whether the system works on live cases as they arrive, used by the people who do the work, against acceptance criteria written in advance. A proof of concept ends in a decision to pilot; a pilot ends in a decision to go live.

Can you skip the proof of concept and go straight to an AI pilot?

Yes, when the task is well understood and you already have a strong evaluation set: a few hundred real cases with expert-approved answers, including the awkward ones. The open question is then whether the system works in your workflow, which a pilot answers. Keep the proof of concept when inputs are unusual, the task is untested on your data or the integrations are expensive.

Where does an MVP fit in an AI project?

An MVP, or minimum viable product, is the smallest version real users can use for real work. In an AI project it usually describes the scope of the pilot: one team, one case type and one channel, with just enough integration and a review screen. Keep the scope minimal but keep the evaluation set, the approval step and the logging, because they show where the system fails.