How to cut LLM costs without quietly cutting quality.
Most savings come from sending less, and sending it to the right place. Every change is checked against the evaluation set, because a cheaper system that needs more review is not cheaper.
veridive6 min read
The model invoice has doubled, and someone asks engineering to switch to a cheaper model. That is usually the last lever to pull, not the first. Most savings come from sending less, and sending it to the right place: shorter context, caching, batching, routing and output limits.
The catch is quality. Every change is checked against the evaluation set, and review time is watched alongside the bill, because a cheaper system that needs more review is not cheaper. The checklist below runs in that order.
Where does LLM spend actually go?
- Measure cost per finished task, by step. Log calls, input and output tokens, retries and the model used for every step: classification, retrieval, drafting, checking. Tag each call with its workflow and step, so a rise in the bill can be traced to the feature that caused it. Spend usually concentrates in one or two steps, and without per-step numbers you optimize the wrong one.
- Separate the stable part of each prompt from the variable part. Instructions and examples repeated on every call behave differently from case-specific text, and the levers for each differ.
- Put review minutes in the same table. A person’s time checking output is often the largest line, so a change that saves tokens and adds review time is a loss. The full arithmetic is in the cost of a token vs. the cost of a mistake, and forecasting before launch is covered in how to forecast LLM running costs.
With that table in hand, the levers are these:
| Lever | What it saves | Trade-off to check |
|---|---|---|
| Context trimming | Input on every call | Losing the passage that mattered |
| Prompt caching | Repeated input on stable prefixes | Prompt structure; stale content after source changes |
| Batch processing | Price on work that isn’t urgent | Hours of delay; handling failed items |
| Routing to smaller models | Model cost on routine cases | Quality on edge cases; misrouted cases |
| Output limits | Output tokens and reading time | Cut-off answers; missing evidence |
| Retry limits | Wasted calls | More cases for people to handle |
How do you trim prompts and retrieved context?
- Send fewer, better passages. Retrieve fewer chunks, re-rank them, remove duplicates, and strip boilerplate such as email signatures, disclaimers and page headers. Pass only the record fields the step uses. Resist sending whole documents just because they fit in the context window: you pay for every token on every call. Less irrelevant text often improves answers too, but the evaluation set decides.
- Tighten instructions. Remove duplicated rules, stale examples and rules for cases the step never sees, and move checks that belong in code into code. For multi-turn and agent steps, carry a short structured state instead of the whole transcript.
When do caching and batching help?
- Put the stable part of the prompt first. Some providers and serving stacks can reuse the processing of an unchanged prompt prefix, which lowers cost and latency; check your provider’s terms. Instructions, examples and reference documents go first, the case-specific text last. Where identical requests recur, a stored answer can be reused too, but only with the same sources and permissions, and it must expire when a source changes. Cached answers are stored data like any other; if they hold personal data, agree retention and access with your data protection officer.
- Batch what isn’t urgent. Back-catalog enrichment, overnight reclassification and weekly report drafts rarely need an answer in seconds. Asynchronous batch processing, where offered, trades speed for price. Design for partial failure, so the items that failed are retried or sent to a person without re-running the whole batch.
Where does routing to smaller models fit?
- Route by difficulty, with thresholds set on the evaluation set. Routine cases go to a smaller model, harder ones to a larger model, uncertain or risky ones to a person. Routing works inside a task too: a small model can classify a message or pull a few fields even when drafting needs a larger one. Some easy cases need no model at all, because a rule or a lookup handles them. Treat each routing change like a model switch, with side-by-side runs and a staged rollout.
What about output length, retries and tool calls?
- Cap and shape the output. Ask for structured fields where output feeds a system, set length limits, and drop explanations nobody reads, while keeping the evidence a reviewer needs.
- Limit retries and fix their cause. A high retry rate is a quality problem that shows up as cost. Find the step that fails, fix its schema or prompt, and cap retries so failures go to a person. Budget agent steps and tool calls the same way, and don’t repeat a lookup already made within the task.
A cheaper system that needs more review is not cheaper.
How do you prove quality didn’t drop?
- Re-run the evaluation set on every change, by case type. Compare quality, validation failures and latency with the baseline. Change one lever at a time, so each saving and each effect on quality can be attributed. A saving that costs quality on an important case type is not a saving.
- Watch review time and overrides after release. Some losses only show in live work, as slightly worse drafts that take longer to fix. Once a saving holds, lower the cost ceiling and its alerts to match, so the next drift is caught early.
Here is an illustrative example in made-up units; they are not prices or benchmarks. A support-drafting workflow costs 10 units per finished case: 4 in model calls and 6 in review time.
| Change | Model | Review | Total per case |
|---|---|---|---|
| Before | 4 | 6 | 10 |
| Trim context, cache the stable prefix | 2.2 | 6 | 8.2 |
| Fix the schema behind most retries | 1.8 | 6 | 7.8 |
| Move drafting to a smaller model | 1.2 | 7.5 | 8.7, so rolled back |
The last change looked best on the invoice but cost more in total than the step before it, because the drafts needed more editing. The first two held on every case type in the evaluation set, and they stayed.
Measure before you cut
Add per-step cost logging and put review minutes next to it for a representative stretch of work; the biggest lever usually becomes obvious. In our services, cost control with the evaluation set as the gate sits in AI reliability, and routing and model choice on your own examples in data and AI foundations.
Ask an assistant about this note