veridive is now an applied AI company. Looking for the answer engine?Looking for the answer engine? What happened

veridive TR Start a project Menu

Field notesDelivery

How to write an AI RFP that gets answers you can compare.

Describe the workflow, real examples and what good looks like instead of a feature list. Ask every bidder how they would evaluate, who would build, what you would own and what it costs to run, then decide with a short paid test.

veridive6 min read

Five proposals arrive for the same AI project. One prices a platform license, one a team for six months, one a fixed fee for “phase one”, and two quote accuracy figures measured on different things. Nobody can compare them, because the RFP listed features and let each bidder imagine the project.

An AI RFP gets comparable answers when it describes the work instead: the workflow, real examples and what good looks like. It asks every bidder the same questions about evaluation, team, ownership and running cost. And it ends with a short paid test on the same examples, because results on your cases compare far better than slides.

Why do AI RFPs get answers you can’t compare?

Three habits borrowed from software procurement do the damage. Feature lists (“must support retrieval”, “must include a chatbot”) invite every bidder to say yes to everything. An undefined target such as “high accuracy” lets each bidder define accuracy on its own data. And without real examples, each bidder pictures different cases, so each prices a different project.

Commercial models make it worse. A license, a team for six months and a fixed fee per phase can’t be compared unless the RFP asks for cost in the same units: a fixed price for the first phase, and a running cost per case at your volumes.

If the RFP doesn’t define the work, each bidder defines it for you.

What should you describe about the work?

The things a bidder would otherwise have to guess:

  • Workflow and volumes: the steps, who does them, the owner, weekly volume and peaks.
  • Systems: what the solution must read from and write to.
  • Data and languages: sources, sensitivity, formats and the languages in play, Turkish and English mixed in one message included.
  • Constraints: where data may be processed, hosting preferences, security reviews, and the KVKK and GDPR questions your data protection officer wants answered.
  • Definition of good: what a correct outcome is, which errors are never acceptable and what should go to a person.
  • Sample cases: a small, masked set of real cases with the answers your experts would accept, awkward ones included.

The sample does more for comparability than any other section. Check with your data protection officer before sharing it: what must be removed, under which agreement it goes out, and when bidders must delete it.

What should bidders be asked to explain?

The same required answers from everyone, in the same order, in a template with a word limit per answer, so you compare substance rather than presentation:

  1. Evaluation plan. How they would build an evaluation set from your cases, who writes the reference answers, and how results are reported by case type against acceptance criteria.
  2. The team who will build. Names, roles, time allocated, and what they need from your side.
  3. Ownership. Who owns the code, prompts, evaluation sets and documentation, and the terms for anything the bidder keeps.
  4. Running-cost model. Cost per case at your volumes, review time included, what drives it up, and how a ceiling is enforced.
  5. Exit. What the handover contains, how you would switch model or supplier, and what happens to your data.
  6. Approval points. Where a person decides, and what they would not automate.

What should you leave out?

  • Model names and versions. They change faster than tenders. Specify the evaluation the system must pass, and ask for a design that can switch models.
  • Accuracy targets without a definition. Ask bidders how they would measure quality on your sample instead.
  • Generic product questionnaires. Security and product due diligence belong in a separate vendor questionnaire.
  • Requests for free proofs of concept. Unpaid work gets rushed work. Pay for a short, bounded test.
  • A fixed price for an undefined program. Ask for a fixed price per phase, with the next phase decided on evidence.

How should proposals be evaluated?

Agree the scoring before the first proposal is opened: the required answers, fit with your constraints, and price, with weights set in advance. Score answers side by side, not proposal by proposal, and ask each shortlisted team to answer the questions that show how they really work in person, with the people who would build in the room.

Then run a short paid test for two finalists on the same sample, against the same acceptance criteria and the same time box. Compare quality by case type, cost per case and how each team worked with yours. The winner’s first phase is then a bounded pilot with the criteria you already wrote, the shape of a Pilot to Production engagement. For what moves the price, see what drives the cost of an AI project.

What does an RFP outline look like?

Here is an illustrative RFP for triaging customer emails at an online retailer, each section filled in briefly.

SectionIllustrative entry
Workflow and volumesEmails to the service inbox are read, tagged by intent and routed to one of six queues. A few thousand a week, peaking after campaigns. Owner: head of customer service.
SystemsReads the shared inbox, CRM and order system. Writes only a tag and a queue in the CRM.
Data and languagesTurkish and English, sometimes both in one email. Names, addresses, order numbers, photo attachments.
ConstraintsProcessing location and retention as agreed with our DPO. Our own cloud preferred. No use of our data for training.
Definition of goodRight queue on the first pass. Safety complaints and legal threats never misrouted. Unclear cases go to a person.
Sample casesA few dozen masked emails with the correct queue for each, shared under agreement after DPO review.
Required answersEvaluation plan, named team, ownership, running cost per email, exit and handover, approval points.
DecisionTwo finalists run a short paid test on the sample.

Before the first draft

Write the definition of good before anything else. If your team can’t agree on it, the RFP isn’t ready, and that disagreement is the most useful thing you have learned so far. If you want an outside view first, a Discovery Sprint produces the baseline, the data and access review and the acceptance criteria an RFP like this needs. Answering tenders rather than writing one? See AI for RFP responses.

Ask an assistant about this note

DeliveryProcurementEvaluation

veridive

Field notes are written and reviewed by veridive. How we write them

Questions

Questions about this note

What should an AI RFP include?

A description of the workflow and its volumes, the systems involved, the data and languages, constraints such as hosting and data protection, and a definition of good that names the unacceptable errors. Add a small masked sample of real cases, and require every bidder to answer the same questions on evaluation, team, ownership, running costs and exit.

Should an AI RFP specify which model to use?

Usually not. Models and prices change faster than procurement cycles, and naming one invites bidders to sell that model rather than solve the workflow. Specify the outcome instead: the evaluation the system must pass on your examples, the hosting and data constraints, and a design that lets you switch models later without rebuilding the system.