How to forecast LLM running costs before the first invoice arrives.
Forecast from the task, not from a price list: calls per task, tokens per call, retries, volume peaks and review minutes, all measured in the pilot. Then set a ceiling with alerts before launch.
veridive6 min read
A pilot’s model bill is usually small enough to ignore, which is exactly why production bills surprise people. The pilot ran on a few hundred friendly cases. Production runs on every case, including the long ones, the retries and the month-end rush.
Good LLM cost estimation starts from the task, not from a price list. Measure calls per task, tokens per call, retries and review minutes in the pilot, project them to full volume and to the peaks, and set a ceiling with alerts before launch. Here are the six steps, each with what “done” looks like and what tends to go wrong.
What goes into the cost of one task?
Forecast from the task, not from the price list.
A task is one finished unit of work: one invoice drafted, one ticket answered. It rarely means one model call. A typical flow classifies the input, retrieves context, drafts, checks the draft and sometimes retries, and each step has its own input and output tokens, as our note on what a token is explains.
So write the flow down with every call in it: the embedding call behind retrieval, the validator that checks the output, the fallback to a larger model, and the instructions resent with each call.
- Done when: every step and every call is listed, with the model each one uses.
- What goes wrong: counting only the main call, and forgetting retrieval, validation and retries.
How do you measure tokens and calls during a pilot?
Log usage from the API responses, not from word counts. Providers report input and output tokens per request; tag each record with the case ID, the step and the language, and keep the retries with their causes.
Measure across the real mix of cases, not a hand-picked set. Tokens per word vary by language and by model, so measure real Turkish and English text rather than assuming a ratio: the same case can take noticeably more tokens in one language than the other. Record review minutes too, from opening a case in the review screen to approving it.
- Done when: you have typical and long-case figures per step and per language, not one average.
- What goes wrong: measuring short demo cases, or a different model from the one you will run. If the choice is still open, compare candidates on the same evaluation set, as in our model selection work.
How do you project to full volume and peaks?
Take monthly volume from the system of record, not from the pilot, and add growth as people come to trust the system. Then find the peaks: month-end close, campaign weeks, the returns season after the holidays. Express each as a peak factor, the busiest period divided by an average one.
Budget on the busiest month, and size capacity and the provider’s rate limits for the busiest day. Check the case mix at the peak as well: month-end invoices and post-holiday returns are often longer and messier than the average case.
- Done when: the forecast has an average month, a peak month and a peak day.
- What goes wrong: treating the pilot month as typical, or forgetting that peaks change the mix as well as the volume.
Which costs sit outside the model bill?
The model bill is one line. The others:
- Review time, often the largest: minutes per case, times volume, times what an hour of that team costs.
- Infrastructure: hosting, the search index, storage, logging and tracing tools.
- Re-indexing when documents change, since every re-embedded page is billed again.
- Evaluation runs: running the evaluation set on every change costs tokens too.
- Monitoring, support and improvement: the people who keep the system useful.
Our note on the cost of a token vs. the cost of a mistake adds the line forecasts often leave out: the expected cost of errors.
- Done when: every line has an owner and a monthly estimate.
- What goes wrong: leaving out review time because “people were already doing this work”.
How do you set a ceiling and alerts?
Agree a monthly ceiling with the owner, based on the peak-month forecast plus a margin. Then add three kinds of alert:
- Budget alerts well before the ceiling, for example at half and at three-quarters of it.
- A daily spend alert, which catches a runaway loop within a day rather than at month-end.
- A cost-per-task alert, which catches prompt bloat or a retry storm even while totals look normal.
Split budgets by feature, so you can see which part caused an increase. Decide in advance what happens near the ceiling: routine cases move to a smaller model, non-urgent work waits in a queue, or cases fall back to people. It should never be a silent stop.
- Done when: each alert has been triggered once on purpose and reached a named person.
- What goes wrong: alerts that land in a shared inbox nobody reads.
What does a forecast sheet look like?
One row per step, with these columns:
- Step, and the model it uses.
- Calls per task, including fallbacks.
- Input and output tokens per call, typical and long.
- Retry rate for the step.
- Monthly volume and peak factor, shared by all rows.
- Review minutes per task, on a line of their own.
Take an illustrative forecast for invoice drafting. The numbers are invented to show the arithmetic, in made-up units; they are not prices or benchmarks. Assume 1 unit per 1,000 input tokens and 5 units per 1,000 output tokens.
| Step | Calls per task | Tokens per call (in / out) | Units per task |
|---|---|---|---|
| Classify | 1 | 2,000 / 100 | 2.5 |
| Extract fields | 1 | 5,000 / 400 | 7 |
| Match and check | 1 | 3,000 / 200 | 4 |
| Draft entry | 1 | 2,000 / 300 | 3.5 |
That is 17 units per invoice. One extraction in ten runs twice, adding 0.7, so 17.7. At 10,000 invoices in an average month the model costs 177,000 units; the busiest month, at a peak factor of 1.5, costs 265,500. Review adds 1.5 minutes per invoice at an invented 10 units per minute: 15 units per invoice, almost as much as the model. The ceiling is set on the peak month, and a per-task alert fires above, say, 25 units.
Cost estimator
Run the forecast with your own numbers.
The fields start with the example above, averaged over its four steps, in its made-up units. Replace them with your pilot’s figures and your provider’s current prices, in any currency. Nothing you type is stored or sent.
The estimator needs JavaScript. The arithmetic it runs is below, so the same forecast works on paper or in a spreadsheet.
Enter a number, 0 or more.
Result
- Cost per task
- units
- Peak month, the line to budget on
- units
| units | Per task | Average month | Peak month |
|---|---|---|---|
| Model | |||
| Review | |||
| Total |
Infrastructure, re-indexing, evaluation runs and support come on top of these two lines.
The arithmetic
- Model cost per task = calls per task × (1 + retry rate) × (input tokens × input price + output tokens × output price) ÷ 1,000,000
- Review cost per task = review minutes ÷ 60 × the cost of an hour of review
- Cost per task = model cost per task + review cost per task
- Average month = cost per task × tasks per month
- Peak month = average month × peak factor
Instrument the pilot
Add usage logging per case, step and language to your pilot before it ends; every other number in the forecast depends on it. Once the system is live, AI reliability keeps it under that ceiling with alerts and routing.
Ask an assistant about this note