What makes an AI pilot ready for production?
A pilot proves something is possible. Production proves it’s reliable, affordable and owned. Our checklist covers evaluation thresholds, approval points, cost ceilings, monitoring and the runbook your team will actually use.
veridive8 min read
A pilot answers a question: can this work? Production makes a promise: it will keep working, day after day and after every model update, at a cost the business has accepted, with a named person responsible when it doesn’t. Plenty of AI pilots answer the first question convincingly and then stall on the second.
The gap is rarely the model. It is everything around it: the criteria nobody wrote down, the approval step nobody designed, the bill nobody forecast, the runbook nobody owns. Before any system we build goes live, we check eight conditions with the person who owns the workflow. If one is missing, the system waits. Here is the list, and why each item is on it.
Why do good pilots stall?
Pilots run under friendly conditions. The examples are hand-picked, the builders sit next to the users, and everyone forgives the odd mistake because it is “just a pilot”. Production removes all three. Inputs arrive messy and incomplete, the builders move on, and a wrong answer now has a customer, an amount or a deadline attached to it.
The stall points are predictable:
- No shared definition of good. The demo impressed, but nobody agreed what accuracy, speed or cost would be acceptable at scale.
- Late reviews. Security, legal and data protection see the system a week before launch, and rightly ask questions that take months to answer.
- Unknown running costs. Pilot usage was small, so nobody projected the bill at full volume.
- No owner. The pilot team leaves, and nobody on the business side knows how to operate, fix or improve the system.
Each of these is cheaper to solve before the pilot than after it. That is why the checklist starts on the first day of a pilot, not in its final week.
Quality: agree what good means, in writing
1. Acceptance criteria agreed and written. Before building, the workflow owner and the team write down what the system must achieve to go live. Good criteria can be tested: the share of cases where the recommendation matches the expert decision, the longest acceptable wait for an answer, the cost per case the business case can carry. They also name the errors that are unacceptable at any rate, such as a wrong refund amount or advice that contradicts policy.
2. Evaluation set built from real examples. Criteria need something to be tested against. We build an evaluation set from real cases, usually a few hundred, each paired with the outcome an expert would accept. The cases come from live work, not a demo script, and they include the awkward ones: incomplete forms, unusual requests, cases where the right answer is to escalate. Part of the set is never used to tune prompts, so the final test stays honest.
Control: decide who sees what, and who decides
3. Data classified and access mapped. We list every source the system reads, what it contains, where it is processed and who may see the output. The system inherits the permissions people already have: it should never show someone a document they could not open themselves. Processing locations and retention are agreed with IT and compliance, in line with KVKK and GDPR requirements.
4. Human approval points defined. For every action the system can take, we ask: can it be undone, does it cost money, does it reach a customer, would we want a second pair of eyes? Where the answer calls for a person, the design includes an approval step, and the reviewer sees the evidence, not just the recommendation. When the system cannot complete a task reliably, the case goes to a person with everything gathered so far.
5. Prompt-injection and data-leak tests passed. An AI system reads text written by other people: emails, documents, forms. Some of it may contain instructions that try to change what the system does. Before launch we test with adversarial inputs, such as a document that asks the system to reveal other customers’ data or skip an approval, and confirm that it refuses. These tests join the evaluation set, so they run again after every change.
A pilot proves something is possible. Production proves it’s reliable, affordable and owned.
Cost: know the bill before it arrives
6. Cost ceiling and alerts configured. During the pilot we measure the full cost per case: model usage, infrastructure and the time people spend reviewing. We project it to production volume and agree a ceiling with the owner. Alerts fire well before the ceiling, and the design says what happens if it is reached: routine cases move to a smaller model, or the queue waits for people, instead of a surprise on the next invoice. Our note on the cost of a token vs. the cost of a mistake shows the arithmetic.
Operations: watch it, and know what to do when it breaks
7. Monitoring and regression checks running. In production, quality can change without anyone touching the system: a provider updates a model, a policy document is replaced, a new product line brings new kinds of cases. We monitor quality through a regular sample reviewed by people, alongside latency, cost and usage. Every change to a model, prompt or data source triggers a re-run of the evaluation set before it goes live. A rising rate of human overrides is one of the most useful early warnings; what to monitor after launch lists the rest.
8. Runbook, documentation and training delivered. The runbook answers the questions someone will ask on a bad day: what to do when quality drops, when a source system is down, when costs spike, when a user reports a harmful answer, how to roll back and who to call. Documentation covers prompts, data flows and the decisions behind them. Users and owners get role-based training, so the people who run the system can also improve it.
Who signs it off?
The owner of the workflow, not the team that built it. In a short go-live review we walk through the eight conditions with evidence for each: evaluation results, the access map, the approval design, the test log, the cost projection, the monitoring view and the runbook. There are three possible outcomes: go live, go live with limits (one team, channel or case type first), or wait until a condition is met. The decision and its reasons are written down.
What does this look like on a real workflow?
Take an illustrative example: a pilot that drafts ERP entries from supplier invoices for a finance team. The acceptance criteria say how often a draft must match the entry an accountant would post, and that a wrong amount or a wrong supplier is never acceptable. The evaluation set holds a few hundred past invoices, including credit notes, handwritten corrections and invoices in two languages. The access map shows that the system reads invoices and purchase orders but never payroll. An accountant approves every entry before it posts, and anything the system is unsure about arrives flagged. The injection tests include an invoice with hidden text asking for the supplier’s bank details to be changed. The cost projection covers month-end peaks, not an average day. And the runbook says what the team does when the ERP is down on the last working day of the month.
None of this is exotic. It is simply written down before launch rather than discovered after it.
What changes after go-live?
The first weeks run with closer review: people check a larger sample of cases, and the owner meets the team weekly. When the numbers hold, review settles into its normal rhythm and the scope widens one step at a time, as in a rollout from one team to many. Two items never finish: cost control and monitoring run for as long as the system does.
The short version for your next steering meeting
Ask these eight questions about any AI pilot heading for production:
- Have we written down what “good enough” means?
- Do we have real examples with expert-approved answers to test against?
- Do we know which data the system reads and who can see its output?
- Do we know where a person approves, and what happens when the system is unsure?
- Have we tried to trick it, and did it hold?
- Do we know the monthly cost at full volume, and what happens at the ceiling?
- Will we notice if quality drops after launch?
- Does someone on our side know how to run it on a bad day?
If the answer to any of them is “not yet”, you have a pilot, not a production system. That is fine, as long as everyone knows which one they are looking at.
Ask an assistant about this note