# How to forecast LLM running costs before the first invoice arrives.

> In this field note, veridive explains how to forecast LLM running costs from the task rather than a price list: calls per task, input and output tokens, retries, volume, peak factors and review minutes, measured in a pilot. It covers costs outside the model bill, ceilings and alerts, and a forecast sheet with an illustrative example.

Forecast from the task, not from a price list: calls per task, tokens per call, retries, volume peaks and review minutes, all measured in the pilot. Then set a ceiling with alerts before launch.

## Key takeaways

- Forecast from the task: every call, its input and output tokens, retries and review minutes, measured in the pilot.
- Measure tokens on real Turkish and English cases, because tokens per word vary by language and model.
- Budget on the peak month and size capacity for the peak day, not for the average.
- Set a ceiling with budget, daily and per-task alerts before launch, and decide what happens near it.

A pilot’s model bill is usually small enough to ignore, which is exactly why production bills surprise people. The pilot ran on a few hundred friendly cases. Production runs on every case, including the long ones, the retries and the month-end rush.

Good LLM cost estimation starts from the task, not from a price list. Measure calls per task, tokens per call, retries and review minutes in the pilot, project them to full volume and to the peaks, and set a ceiling with alerts before launch. Here are the six steps, each with what “done” looks like and what tends to go wrong.

## What goes into the cost of one task?

> Forecast from the task, not from the price list.

A task is one finished unit of work: one invoice drafted, one ticket answered. It rarely means one model call. A typical flow classifies the input, retrieves context, drafts, checks the draft and sometimes retries, and each step has its own input and output tokens, as our note on [what a token is](https://veridive.com/insights/what-is-a-token-in-ai/) explains.

So write the flow down with every call in it: the embedding call behind retrieval, the validator that checks the output, the fallback to a larger model, and the instructions resent with each call.

- **Done when:** every step and every call is listed, with the model each one uses.
- **What goes wrong:** counting only the main call, and forgetting retrieval, validation and retries.

## How do you measure tokens and calls during a pilot?

Log usage from the API responses, not from word counts. Providers report input and output tokens per request; tag each record with the case ID, the step and the language, and keep the retries with their causes.

Measure across the real mix of cases, not a hand-picked set. Tokens per word vary by language and by model, so measure real Turkish and English text rather than assuming a ratio: the same case can take noticeably more tokens in one language than the other. Record review minutes too, from opening a case in the review screen to approving it.

- **Done when:** you have typical and long-case figures per step and per language, not one average.
- **What goes wrong:** measuring short demo cases, or a different model from the one you will run. If the choice is still open, compare candidates on the same evaluation set, as in our [model selection work](https://veridive.com/services/data-ai-foundations/).

## How do you project to full volume and peaks?

Take monthly volume from the system of record, not from the pilot, and add growth as people come to trust the system. Then find the peaks: month-end close, campaign weeks, the returns season after the holidays. Express each as a peak factor, the busiest period divided by an average one.

Budget on the busiest month, and size capacity and the provider’s rate limits for the busiest day. Check the case mix at the peak as well: month-end invoices and post-holiday returns are often longer and messier than the average case.

- **Done when:** the forecast has an average month, a peak month and a peak day.
- **What goes wrong:** treating the pilot month as typical, or forgetting that peaks change the mix as well as the volume.

## Which costs sit outside the model bill?

The model bill is one line. The others:

- **Review time**, often the largest: minutes per case, times volume, times what an hour of that team costs.
- **Infrastructure**: hosting, the search index, storage, logging and tracing tools.
- **Re-indexing** when documents change, since every re-embedded page is billed again.
- **Evaluation runs**: running the evaluation set on every change costs tokens too.
- **Monitoring, support and improvement**: the people who keep the system useful.

Our note on [the cost of a token vs. the cost of a mistake](https://veridive.com/insights/cost-of-a-token-vs-cost-of-a-mistake/) adds the line forecasts often leave out: the expected cost of errors.

- **Done when:** every line has an owner and a monthly estimate.
- **What goes wrong:** leaving out review time because “people were already doing this work”.

## How do you set a ceiling and alerts?

Agree a monthly ceiling with the owner, based on the peak-month forecast plus a margin. Then add three kinds of alert:

- **Budget alerts** well before the ceiling, for example at half and at three-quarters of it.
- **A daily spend alert**, which catches a runaway loop within a day rather than at month-end.
- **A cost-per-task alert**, which catches prompt bloat or a retry storm even while totals look normal.

Split budgets by feature, so you can see which part caused an increase. Decide in advance what happens near the ceiling: routine cases move to a smaller model, non-urgent work waits in a queue, or cases fall back to people. It should never be a silent stop.

- **Done when:** each alert has been triggered once on purpose and reached a named person.
- **What goes wrong:** alerts that land in a shared inbox nobody reads.

## What does a forecast sheet look like?

One row per step, with these columns:

- **Step**, and the model it uses.
- **Calls per task**, including fallbacks.
- **Input and output tokens per call**, typical and long.
- **Retry rate** for the step.
- **Monthly volume** and **peak factor**, shared by all rows.
- **Review minutes** per task, on a line of their own.

Take an illustrative forecast for invoice drafting. The numbers are invented to show the arithmetic, in made-up units; they are not prices or benchmarks. Assume 1 unit per 1,000 input tokens and 5 units per 1,000 output tokens.

| Step | Calls per task | Tokens per call (in / out) | Units per task |
|---|---|---|---|
| Classify | 1 | 2,000 / 100 | 2.5 |
| Extract fields | 1 | 5,000 / 400 | 7 |
| Match and check | 1 | 3,000 / 200 | 4 |
| Draft entry | 1 | 2,000 / 300 | 3.5 |

That is 17 units per invoice. One extraction in ten runs twice, adding 0.7, so 17.7. At 10,000 invoices in an average month the model costs 177,000 units; the busiest month, at a peak factor of 1.5, costs 265,500. Review adds 1.5 minutes per invoice at an invented 10 units per minute: 15 units per invoice, almost as much as the model. The ceiling is set on the peak month, and a per-task alert fires above, say, 25 units.

## Instrument the pilot

Add usage logging per case, step and language to your pilot before it ends; every other number in the forecast depends on it. Once the system is live, [AI reliability](https://veridive.com/services/ai-reliability/) keeps it under that ceiling with alerts and routing.

## Frequently asked questions

### How do you estimate LLM API costs before launch?

Measure a pilot on real cases: log calls per task, input and output tokens per call, retries and review minutes, by step and by language. Multiply by monthly volume, apply a peak factor for the busiest periods, add costs outside the model bill such as infrastructure and review time, then set a ceiling with alerts before launch.

### Why is the average month the wrong basis for an AI budget?

Because volume and case mix peak. Month-end closes, campaigns and seasonal returns bring more cases, and often longer ones, in a short period, so a budget built on the average month is exceeded at the first peak. Forecast an average month, a peak month and a peak day, budget on the peak month, and check that capacity and rate limits cover the peak day.
