What drives the cost of an AI project, and how to keep it predictable.
The model is rarely the expensive part. Cost follows the systems you connect, the state of the data, the review design, the languages, the security reviews and how much routines change. A fixed scope per phase keeps it predictable.
veridive6 min read
Three suppliers quote for the same AI project, and the prices are so far apart that they seem to describe different projects. Often they do. One priced a prototype, one priced a pilot on one system, and one priced production across the company.
The model is rarely the expensive part. AI implementation cost follows the work around the model: the systems to connect, the state of the data, the review design, the languages, the security and data-protection reviews, and how much the team’s routine has to change. A fixed scope per phase is what keeps that cost predictable.
Why is “how much does AI cost?” hard to answer in one number?
Because “AI project” covers everything from a one-day prototype to a system that reads from three ERPs, writes into one and works in two languages. The model call is similar in both. Almost everything else differs.
It is also hard because the biggest unknowns are found only by looking: whether the data is reachable, how messy the real cases are, how many exceptions the workflow has, and what the security review will ask. An honest supplier can’t price those before someone has looked, which is why the first step should be small, fixed and designed to remove unknowns.
Which build costs matter most?
These seven drivers usually explain the gap between two quotes:
| Driver | Why it matters | How to reduce it |
|---|---|---|
| Integrations | Each system needs access, mapping, error handling and tests; writing costs more than reading | Start read-only, with one system |
| Data access | Approvals, permissions, formats and cleaning come before any AI work | Pick a workflow whose data is already reachable |
| Languages | Each language needs its own examples, reviewers and tests | Launch in one language; add the next with its own cases |
| Review screens | People need the evidence in view to approve in one step | Reuse the tools people already work in |
| Evaluation | Experts write reference answers for a few hundred cases | Budget their time up front |
| Security and data-protection review | Late questions cause rework | Involve reviewers in the first week |
| Change management | Training, new routines and ownership | Start with one team and its real tasks |
The last two are easy to leave out of a quote, and they decide whether the system goes live and gets used.
The model is rarely the expensive part. The work around it is.
What makes a project more expensive than it looks?
A few multipliers hide behind a simple description:
- Writing, not just reading. Posting into an ERP needs approvals, rollback and far more testing than drafting a suggestion.
- Many variants. Every supplier layout, document template or channel adds cases to test.
- Rare cases that must be right. Getting the unusual cases right costs more than getting the common ones right.
- Unclear ownership. Decisions wait, and waiting is billed.
- Scope that grows during the build. “While we’re at it” is the most expensive phrase in a project.
An illustrative case makes the point: the same invoice workflow, scoped twice. In the first scope, invoices for one company are read, matched to purchase orders and drafted as entries in one ERP, in Turkish. In the second, the same work covers three companies on three ERPs, for example SAP, Logo and Microsoft Dynamics, in Turkish and English. The extraction and matching logic is largely shared, so the second scope doesn’t cost three times as much. But it has three integrations, each with its own data model, test environment and approval flow; three finance teams to train; and an evaluation set that must cover every combination of system and language, six in all, because an invoice that works in one combination can fail in another. It costs more to build and far more to test.
What makes it cheaper?
Mostly the opposite choices, made on purpose: one workflow, one owner, one system and one language to start; data already reachable in approved systems; drafts before actions; reviews in the first week; and acceptance criteria written before the build, so tuning stops when the system is good enough. A smaller model can lower running costs where the evaluation set shows it is good enough, but it rarely changes the build cost much.
How do running costs compare with build costs?
Build cost is paid once per phase. Running costs recur every month for as long as the system is used: model usage, infrastructure, the minutes people spend reviewing output, monitoring, and the steady work of improving prompts, sources and the evaluation set. Over the life of a system they can exceed the build cost, and review time is often the largest line.
Measure them during the pilot rather than guessing at the start. Our note on forecasting LLM running costs shows how, and the cost of a token vs. the cost of a mistake adds the price of errors.
How do you get a price you can rely on?
Buy the project in phases, each with a fixed scope and a fixed price, where each phase removes the unknowns that make the next one hard to price:
- Executive Build Day (1 day): a working prototype on one real workflow with the leadership team, and three prioritized opportunities.
- Discovery Sprint (about two weeks): the baseline, the data and access review, the risks, and a pilot plan with acceptance criteria.
- Pilot to Production (about six to ten weeks): a bounded pilot on approved data. When it meets the agreed criteria, a separate production scope follows.
These are the formats we offer, and a pilot is how our custom AI software work usually starts. Whoever you work with, ask every supplier the same questions so the quotes can be compared:
- What exactly is in scope: which workflow, systems, languages and case types?
- What is out of scope, and what happens if something new turns up mid-way?
- Which acceptance criteria define “done”, and who tests them?
- What do you assume about our data and access?
- How much of our people’s time do you need, and whose?
- What will it cost to run each month at our volume, and on what basis?
- What do we own at the end: code, prompts, evaluation sets, documentation?
- Is the price fixed for the phase, and what counts as a change?
Make the quotes comparable
Write one paragraph per candidate workflow: the systems it touches, the languages it runs in, its volume and its owner. That is most of what any supplier needs for a first estimate, and it makes the answers comparable. If you’d like a fixed-scope recommendation in writing, describe the workflow.
Ask an assistant about this note