# How to choose an AI partner: questions that show how they really work.

> In this field note, veridive sets out twelve questions for choosing an AI consulting partner, grouped by delivery, people, evaluation and ownership, with why each matters and what a good answer sounds like. It covers red flags such as accuracy promises made before seeing your data, and how to compare firms fairly with a short paid test.

The answers that matter are about evidence and ownership: whether a firm can show an evaluation set from past work, who will actually build, what it refuses to automate, how it estimates running costs and what you own when it leaves.

## Key takeaways

- Judge an AI firm on evidence and ownership, not the pitch: evaluation, builders, limits, running costs and handover.
- Meet the people who will build before you sign, and ask how much of your owner’s time they need.
- Treat accuracy promises made before anyone has seen your data, demos on their own examples and mandatory platforms as red flags.
- Compare firms with the same brief, the same questions and a short paid test on your own examples.

Every AI proposal promises an experienced team, a proven method and responsible AI. The slides look alike, the demos all impress, and the prices are hard to compare because each firm scoped a different project.

The differences show up in the answers to a few specific questions, and nearly all of them are about evidence and ownership. Can the firm show how it measured quality on past work? Who will actually build? What would it refuse to automate? How does it estimate running costs? What will you own when it leaves? A firm that answers these plainly is easier to work with than one that answers them beautifully.

## What should you look for before the first meeting?

Start with your side. Pick one workflow, name its owner, gather a handful of real examples you are allowed to share, masked if needed, and note roughly how the work runs today. A firm’s response to your real cases tells you more than any capability deck. If you are still deciding whether to [build, buy or wait](https://veridive.com/insights/build-vs-buy-ai/), settle that first: a services firm is the answer to only one of the three.

Then read what each firm publishes. Look for specifics: named workflows, what was measured against which baseline, illustrative examples labeled as such, and a stated list of things it won’t do. Check whether there is a small, fixed-scope first step you can buy without committing to a program.

## Which questions reveal how a partner builds, and who does the work?

**Delivery**

1. **Which workflow would you start with, and what would you measure first?** Good answers start from work, not technology: one workflow, an owner and a baseline of volume, time per case, errors and cost.
2. **Where would a person approve, and what would you refuse to automate?** This tests judgment. Listen for specific decisions that stay with people, and at least one thing the firm would not automate in your case.
3. **How do you estimate running costs?** Good answers price a finished task, review time included, project it to full volume and set a ceiling with alerts.

**People**

4. **Who exactly will build this, and can we meet them?** The pitch team and the build team are often different people. Good answers name them and offer the meeting before you sign.
5. **How much of their time is ours?** You want a stated allocation, and a plan for when someone leaves.
6. **What do you need from our side?** A good firm asks for your workflow owner’s time and for experts to write reference answers. A firm that needs nothing from you is planning to guess.

## Which questions reveal how they measure quality?

7. **Can you show an evaluation set from past work?** Client data may be confidential, but the structure isn’t: real cases, reference answers written by experts, case types, scoring rules, results by case type. A firm that can’t show the format has probably never built one.
8. **How would you build one for our workflow?** Listen for your real cases, awkward ones included, reference answers written by your people rather than theirs, and a portion held back for the final test.
9. **What would make you recommend stopping?** Acceptance criteria only mean something if missing them has a consequence. Good answers name the conditions.

## Which questions reveal who owns what?

10. **Who owns the code, prompts, evaluation sets and documentation?** The good answer is you, in writing. Ask counsel to check the contract wording.
11. **What does the handover include, and how would we test it?** Listen for a runbook, training and a drill in which your team changes something without the firm. The [handover checklist](https://veridive.com/insights/ai-project-handover-checklist/) lists what to expect.
12. **Which models and platforms would we depend on, and how would we switch?** Good answers keep the model replaceable, deploy where your data needs to live and include an exit plan.

## What are the red flags?

- **Accuracy promises before seeing your data.** No test on your cases means no basis for the number.
- **Demos only on their own examples.** Theirs were chosen to impress; yours were not.
- **A platform you must adopt** before anyone has looked at the workflow.
- **No answer on running costs,** or only a price per token.
- **A pitch team you never see again,** and no questions about who owns the work on your side.

> An accuracy figure quoted before anyone has seen your data is a guess.

## How do you run a fair comparison?

Give every firm the same brief, the same examples and the same questions, and score the answers side by side. Then fund a short paid test for the two strongest, on the same examples and against the same written criteria. Comparing slides compares writing; comparing results on your cases compares work. The [RFP template](https://veridive.com/insights/ai-rfp-template/) shows how to structure the brief so the answers line up.

Here are two illustrative proposals for the same workflow, drafting replies to supplier emails.

| Question | Proposal A | Proposal B |
|---|---|---|
| How is quality shown? | A fixed accuracy figure on page one | A paid test on a few hundred of your past emails |
| Who builds? | “A dedicated team” | Three named people, meeting offered |
| Where does a person approve? | Full automation as the goal | Your team approves drafts; nothing is sent automatically |
| What does it cost to run? | Included in a platform fee, up to a volume cap | Cost per email measured in the test, projected to full volume |
| What do you own? | Configuration lives on their platform | Code, prompts and evaluation set are yours |

Proposal A sounds more certain. Proposal B is the one you can check: its uncertainty is honest and bounded by a test you control, while A’s certainty hasn’t met your data yet.

## Compare like with like

Put the twelve questions to every firm you are considering, including us. Ask veridive for the evaluation plan, the names of the people who would build and the ownership terms in writing, and compare them with everyone else’s. [How we work](https://veridive.com/approach/) and [our formats](https://veridive.com/services/) are published, and a [first conversation](https://veridive.com/contact/) is followed by a written recommendation with scope and cost that you can hold next to the others.

## Frequently asked questions

### What questions should you ask an AI consulting firm?

Ask how they would measure quality on your workflow, and whether they can show the structure of an evaluation set from past work. Ask who exactly will build and whether you can meet them, what they would refuse to automate, how they estimate running costs, and who owns the code, prompts and evaluation sets when the project ends.

### What are the red flags when choosing an AI partner?

An accuracy figure promised before anyone has seen your data, demos that only ever use the firm’s own examples, a platform you must adopt before the workflow is understood, and no clear answer on running costs. Others include a pitch team you never see again, no questions about who owns the workflow on your side, and full automation with no approval points.

### How do you compare AI consulting proposals fairly?

Send every firm the same brief with the same real examples, ask the same questions, and score the answers side by side. Then fund a short paid test for the strongest candidates on the same examples, against the same written acceptance criteria. Meet the people who would build, and check in writing what you would own at the end.
