veridive is now an applied AI company. Looking for the answer engine?Looking for the answer engine? What happened

veridive TR Start a project Menu

Field notesEconomics

The cost of a token vs. the cost of a mistake.

Model prices keep falling; mistakes don’t get cheaper. We compare cost per task with the cost of an error — and often end up with a small model for most cases and a person for the rest.

veridive6 min read

Model prices have fallen fast, and they keep falling. That is good news, and it has created a habit: comparing AI options by the price of a token, or by the cheapest model that seems to work. The price of a token matters. But in most business workflows it is the smaller number. The bigger one is the cost of a mistake.

A finished task has two prices: what it costs to do, and what it costs when it goes wrong. A wrong refund costs the same whether a large model or a small one approved it.

What does a task really cost?

Start by pricing a finished task, not a model call; forecasting running costs shows how to measure one. Cost per task includes:

  • Model usage: the tokens for every call the task needs, including retries and the retrieval steps around it.
  • Infrastructure: hosting, storage, search indexes and, for on-premises deployments, the servers themselves.
  • Human review: the minutes a person spends checking, correcting or approving the output. This is often the largest line.
  • Build and upkeep: a fair share of the cost of building, integrating, monitoring and improving the system, spread over the tasks it handles.

Measured this way, the gap between a cheap model and an expensive one is sometimes a rounding error next to review time. A slightly more expensive model whose drafts need less correction can be the cheaper system.

What does a mistake cost?

Next, price the errors, together with the person who owns the workflow. Mistakes are not all equal, so list the kinds that matter:

  • Rework: someone notices and fixes it. The cost is their time.
  • Direct loss: a refund paid that shouldn’t have been, a discount applied twice, a duplicate order shipped.
  • Customer impact: a wrong answer that brings a second contact, a complaint or a lost customer.
  • Compliance exposure: advice that contradicts policy or regulation, or data shown to the wrong person.
  • Lost trust: people stop using the system after it embarrasses them, and the investment quietly stops paying back.

Rough ranges are enough. The point is not precision. It is to put the cost of an error on the table in the same conversation as the cost of a token.

The cheapest model is not always the cheapest system.

A worked example

Take an illustrative workflow. The numbers are invented to show the arithmetic, in made-up units; they are not prices or benchmarks.

Every case gets a quick human check costing 2 units, and a mistake that slips through costs 50 units to put right. Seventy in every hundred cases are routine; thirty are hard. We compare three designs on the same evaluation set.

  • Small model only. Model cost 0.1 per case. It makes 1 mistake on the routine cases and 9 on the hard ones: 10 in every 100. Expected cost per case: 0.1 for the model, 5 for errors, 2 for review, so 7.1 units.
  • Large model only. Model cost 1 per case. It makes 1 mistake on the routine cases and 3 on the hard ones: 4 in every 100. Expected cost: 1 + 2 + 2 = 5 units.
  • Routing. The small model takes the routine cases and the large model takes the hard ones. Mistakes stay at 4 in every 100, but the average model cost falls to 0.37. Expected cost: about 4.4 units.

In this example, the model with the cheapest tokens produces the most expensive system, and routing beats both single-model designs. One more step often helps: send the cases where even the large model is unsure to a person, with the evidence prepared. If those cases hold most of the remaining mistakes, a few minutes of review costs less than the errors it prevents.

Where do the numbers come from?

Not from a vendor’s price list alone. Error rates come from the evaluation set: real cases from your workflow with expert-approved answers, run through each design. Cost per task comes from measuring a pilot, including the minutes people actually spend reviewing. The cost of each kind of mistake comes from the workflow owner and finance, using what similar errors cost today. None of these needs to be perfect. They need to be written down, so the business case can be checked and updated.

What does today’s process cost?

The fair comparison is not an AI design against zero. It is an AI design against the way the work runs now. The current process has its own cost per case, mostly people’s time, and its own error rate, which is rarely zero and often unmeasured. That is why we record the baseline before building. An AI design earns its place when its total cost, mistakes included, beats the baseline by enough to justify the change.

Why does routing usually win?

Routing sends each case to the least expensive option that can handle it reliably. It works because cases are not equally hard: most workflows have a large body of routine work and a smaller set of difficult or risky cases.

  • Routine, low-risk cases go to a small, fast model.
  • Harder cases, or ones where the small model signals low confidence, go to a larger model.
  • High-risk cases, and those where both are unsure, go to a person, with the work prepared.

The thresholds are set on the evaluation set, not by instinct, and checked again after launch. This is how we often end up with a small model for most cases and a person for the rest.

Keep the comparison current

The inputs change. Prices fall, volumes grow, new models appear, and the cost of an error can rise when a workflow reaches more customers. Re-run the comparison when any of them moves, and at least every quarter. When the bill rises, cut costs without cutting quality. Track cost per task and error rates side by side, set a monthly ceiling with alerts, and treat a jump in either as a reason to look again.

Questions to ask of any AI business case

  • What does one finished task cost, including human review?
  • What does each kind of mistake cost, and who pays for it?
  • What is the error rate on our own examples, not on a benchmark?
  • Which cases could a smaller model handle, and which need a person?
  • When will we run the numbers again?

A business case that answers these is worth approving. One that quotes only the price per token is not finished.

Ask an assistant about this note

EconomicsModelsRouting

veridive

Field notes are written and reviewed by veridive. How we write them

Questions

Questions about this note

How do you calculate the cost of an AI workflow?

Add up the full cost per task, not just tokens: model usage, infrastructure, the time people spend reviewing, and a share of build and monitoring costs. Then add the expected cost of errors, which is the error rate multiplied by what a typical mistake costs. Compare that total with what the same work costs today.

Is a cheaper AI model always cheaper to run?

No. A cheaper model that makes more consequential mistakes can cost more in total, once rework, refunds and extra reviews are counted. Often the best result is a routing design: a small model handles routine cases, a larger model handles harder ones, and a person takes the uncertain or high-risk cases.