veridive is now an applied AI company. Looking for the answer engine?Looking for the answer engine? What happened

veridive TR Start a project Menu

Field notesReliability

How to switch language models without breaking what works.

Models get retired, prices change and better options appear, so plan to switch from day one. With an evaluation set, a model-agnostic layer and side-by-side runs, a migration becomes a test result, not a leap of faith.

veridive6 min read

The notice from the model provider is short and polite. The model your customer-email assistant runs on will be retired, and requests to it will stop working after a stated date. The suggested replacement is described as better. Nobody can yet say whether it is better on your cases.

That notice will come, and so will price changes and better options. Plan to switch from day one. With an evaluation set, a model-agnostic layer and side-by-side runs, a migration becomes a test result, not a leap of faith.

Why will you switch models more than once?

Because the reasons to switch come from outside, on other people’s schedules:

  • Retirement notices. Providers retire older models. The notice period is whatever the terms say, not whatever your roadmap needs.
  • Price changes. A cheaper model may now meet your thresholds, or the one you use may cost more.
  • Quality gains. A newer or smaller model may handle your hard case types, or your Turkish text, better on your own examples.
  • Data residency. A requirement to process data in a region, in your own cloud or on-premises can mean a different provider or an open-weight model.

A system that runs for years will meet several of these. Treat switching as routine maintenance, and design for it.

What makes a system easy or hard to move?

Four things make it easy:

  • A model-agnostic layer. Application code calls your own interface, such as “draft reply” or “extract invoice”, and configuration maps each task to a model, a prompt and settings.
  • Prompts stored per model. Each task keeps a versioned prompt for each model, so a new model gets its own tuned prompt without disturbing the old one.
  • An evaluation set with a baseline. You know what “as good as now” means, case type by case type.
  • Validation in code. Outputs are checked against a schema, so a format difference surfaces as a validation failure, not as bad data in the ERP.

What makes it hard: provider-specific features woven through the code, prompts tuned by trial and error with no record of why, parsing that depends on one model’s phrasing, and no baseline. Done well, switching a task to another model is a configuration change plus a test run.

How do you compare the new model with the old one?

Run the candidates side by side on the evaluation set, with the current model as the baseline: first with the existing prompts, then with prompts adjusted for each candidate. Compare quality, validation failures, cost per task and slow-end latency, broken down by case type and language.

Then read the disagreements. Sample the cases where old and new models differ and review them with the workflow owner. Sometimes the new model is right and the reference answer is out of date; sometimes a case type that matters has quietly got worse while the average improved. The comparison is done when there is a table per case type and a decision the owner signs.

With an evaluation set, a migration is a test result, not a leap of faith.

Why do prompts need to change with the model?

Because models read the same words differently. One follows format instructions more literally, another writes longer answers, a third treats examples as templates to copy. Refusals, tone and the handling of ambiguous cases shift too.

The changes don’t stop at the prompt. Output parsing may break on different date formats, field order or escaping; tool-call conventions differ; latency, rate limits and context size change, so timeouts and batch sizes need adjusting. Budget engineering time for all of this in every migration; it is rarely zero. Store each adjusted prompt as a new version tied to its evaluation run, as described in versioning prompts like code.

How do you roll out a new model safely?

In stages, each with an exit criterion written before it starts:

  1. Offline. The evaluation set passes the agreed thresholds on every important case type.
  2. Shadow mode. The new model runs on live traffic in parallel; its outputs are logged and compared, never used.
  3. A gradual share. One queue, channel or case type moves first, while you watch the override rate and validation failures; then the next.
  4. Full switch, with a way back. The old configuration stays ready until the old model is actually retired, so rollback is a configuration change.

During the gradual share, watch the same signals you use every day, described in what to monitor after an AI system goes live.

Go back to the retirement notice above, as an illustrative case. The team runs the evaluation set on two replacements: the provider’s suggested successor and an open-weight model in the company’s own cloud. The open-weight model matches on quality but misses the latency threshold on long threads. The successor passes everywhere except complaint emails in Turkish, where its drafts are stiffer than the tone guide allows. The team adds two real examples of the right tone to that one prompt, re-runs the set and passes. Shadow mode on live email confirms the result, and traffic moves one queue at a time, with the old configuration kept ready until the retirement date.

What should contracts and architecture say about retired models?

Take these questions to procurement and counsel before you sign, because the answers vary by provider and change over time:

  • How much notice is given before a model is retired, and is it written into the agreement?
  • Can you pin a specific model version, and are changes behind a model name announced?
  • Where is data processed, and what changes if you move region or provider? Your data protection officer should see the answer.
  • What can you take with you when you leave: logs, evaluation results, any fine-tuned model?

In the architecture, keep the layer, the prompts and the evaluation set as your own assets, name a tested fallback model for every critical task, and record the model identifier in every log line.

Rehearse the switch

Pick your most important AI task and do a dry run: point it at a second model through configuration and run the evaluation set. If that takes an afternoon, you are ready for the next retirement notice. If it takes a project, do that project before the notice arrives. AI reliability keeps live systems ready to switch; the model-agnostic layer and evaluation set that make it possible are built in data and AI foundations.

Ask an assistant about this note

ReliabilityModelsEvaluation

veridive

Field notes are written and reviewed by veridive. How we write them

Questions

Questions about this note

How do you switch LLM providers safely?

Run the candidate models side by side on your evaluation set and compare quality, cost and latency by case type. Adjust prompts and output parsing for the new model, then run it in shadow mode on live traffic before moving a small share of real work to it, expanding in stages while you watch override rates. Keep the old configuration ready for rollback.

How do you avoid LLM vendor lock-in?

Put a model-agnostic layer between your application and any provider, so each task maps to a model, a prompt and settings in configuration. Store prompts per model, keep your own evaluation set with a baseline, and validate outputs in code. Avoid depending on provider-specific features you can’t take with you, and ask about retirement notice periods before signing.

What happens when an AI model is retired?

Requests to a retired model stop working after the date the provider announces, so any system built on it has to move to another model before then. With an evaluation set and a model-agnostic layer, that move is a planned test and a staged rollout. Without them, it becomes an unplanned project against the provider’s deadline.