# LLMOps and MLOps: what changes when the model is a language model?

> In this field note, veridive compares LLMOps and MLOps. Versioning, testing, monitoring and rollback carry over, but language models are rarely trained in-house: teams manage prompts, retrieval sources, external model versions and human review instead. The note covers what gets versioned, how monitoring differs, who does the work, and what a data team should learn first.

Much of MLOps carries over, but with language models you rarely train: you manage prompts, retrieval sources, external model versions and human review. The center of gravity moves from training pipelines to evaluation and change control.

## Key takeaways

- Versioning, testing, staged deployment, monitoring and rollback carry over from MLOps; retraining mostly doesn’t.
- In an LLM system you version prompts, sources and the index, model identifiers, evaluation sets and configuration together.
- The provider can change or retire the model, so evaluation and change control replace the training pipeline as the center of gravity.
- Human review is an operational component with its own metrics: throughput, queue age, agreement and override rates.

A data team runs a demand-forecasting model with a mature routine: feature pipelines, scheduled retraining, a model registry, drift monitoring and a rollback button. Then the business asks it to operate an assistant built on a language model. Most of the routine still applies. But the first time quality drops, nobody has retrained anything, and the model isn’t even theirs.

Much of MLOps carries over: versioning, testing, monitoring, rollback. What changes is where the work sits. With language models you rarely train; you manage prompts, retrieval sources, external model versions and human review. The center of gravity moves from training pipelines to evaluation and change control.

## What does MLOps cover, and what carries over?

MLOps is the practice of taking a trained model from experiment to production and keeping it useful: data and feature pipelines, experiment tracking, training and retraining, a model registry, deployment, monitoring for drift and accuracy, and rollback.

The habits carry over almost intact. Version everything that affects behavior. Test before release against data the system hasn’t seen. Deploy in stages. Monitor in production against a baseline. Roll back fast. Keep an audit trail. The engineering underneath carries over too: automated pipelines, infrastructure as code, access control.

## What is new with language models?

Five things change the daily work:

- **You rarely train.** Behavior is shaped by prompts, examples, retrieved sources and configuration, and it can change in minutes. Fine-tuning exists, but it is the exception.
- **The model is external.** A provider hosts it, may update it and will eventually retire it. Your team didn’t train it and can’t inspect its training data.
- **Output is language.** There is rarely one accuracy metric; quality is judged with evaluation sets and rubrics, partly by people.
- **People are in the loop.** Reviewers approve and correct output in production, so their work is part of the running system.
- **Cost is per call.** Spend moves with volume, prompt length and model choice, unless you host the model yourself.

| Aspect | MLOps: a model you trained | LLMOps: a language model you use |
|---|---|---|
| Main artifacts | Training data, features, model weights | Prompts, sources and index, model identifiers, evaluation sets, configuration |
| Testing | Metrics on held-out labeled data | Evaluation sets with rubrics, repeated runs, adversarial cases |
| Monitoring | Data drift, prediction drift, accuracy when outcomes arrive | Override rate, sampled review, citation support, cost per task |
| Change triggers | New data, drift, scheduled retraining | Prompt edits, source updates, provider updates and retirements |
| Roles | Data scientists, ML engineers | Engineers, domain owners, content owners, reviewers |

Many real systems need both. A trained model scores, forecasts or classifies, and a language model reads and drafts around it. Where you train, MLOps applies unchanged; where you call a language model, the rest of this note applies.

## What gets versioned in an LLM system?

Everything that can change an answer:

- **Prompts,** per task and per model, reviewed like code; see [versioning prompts like code](https://veridive.com/insights/prompt-versioning-for-production/).
- **Sources and the retrieval index:** which documents, which versions, and the chunking and embedding settings, so the index can be rebuilt exactly.
- **Model identifiers:** the exact version string for each task, not just a family name, recorded in every log line.
- **Evaluation sets,** so scores from different releases can be compared.
- **Configuration:** sampling settings, length limits, routing thresholds, tool lists and permissions.

A release is the combination of all five. Any answer in production should be traceable to one version of each.

> Your team didn’t train the model, but it still owns every change to its behavior.

## How does monitoring differ?

A forecasting model eventually meets the truth: actual demand arrives, and accuracy can be computed. A language model’s answers rarely come with ground truth. The review step provides it through overrides, edits and sampled review scores, alongside citation support, validation failures, cost per task and latency. Drift has new sources too, including a provider’s model update when your inputs haven’t changed at all; [what to monitor after an AI system goes live](https://veridive.com/insights/llm-monitoring-metrics/) covers the signals.

Human review needs monitoring of its own, as an operational component: throughput, queue age, agreement between reviewers, and override rates by reviewer, where outliers point to training needs or to reviewers who have stopped checking. It needs staffing and cover like any other part of the system.

## Who does the work in each?

In MLOps, data scientists build and retrain models, ML engineers deploy and monitor them, and data engineers run the pipelines. In an LLM system, software engineers own the application, prompts and integrations; data engineers own the sources and the index; domain experts own the evaluation set; content owners keep documents current; reviewers work the queue; and the workflow owner sets thresholds and approves changes.

The biggest shift is that people outside the data team now change the system’s behavior, for example by replacing a policy document. Change control has to reach them.

Take an illustrative change request: the company launches a new product line.

- **In the machine learning pipeline** for demand forecasting, the team gathers history for similar products, builds features, retrains, validates on a held-out period and deploys. Until real sales arrive, forecasts for the new line rest on proxies and are flagged. Most of the work is data and training.
- **In the LLM system** for sales support, nobody retrains anything. The content owner adds the catalog and price list to the sources, and the index is rebuilt. The prompt gets one rule about the line’s warranty terms. The domain owner adds a few dozen real questions about the line to the evaluation set, the full set runs to confirm nothing else moved, and the release rolls out in stages while the override rate for the new case type is watched. Most of the work is content, evaluation and change control.

## What should a data team learn first?

Evaluation first: building evaluation sets with domain owners, writing rubrics, reading results by case type. Then retrieval, which is natural data-team territory: parsing, chunking, search and permissions. Then change control for things that aren’t code, such as prompts and documents. Training models from scratch can wait; most LLM systems never need it.

## Version everything that matters

List every artifact that can change your system’s answers, from prompts to the documents in the index, and check that each is versioned and tied to an evaluation run. Where one isn’t, that is the first gap to close. This is how [AI reliability](https://veridive.com/services/ai-reliability/) operates live systems, on sources, indexes and evaluation sets built in [data and AI foundations](https://veridive.com/services/data-ai-foundations/).

## Frequently asked questions

### What is LLMOps?

LLMOps is the set of practices for running applications built on large language models in production. It covers versioning prompts, retrieval sources, model identifiers, evaluation sets and configuration; testing changes against an evaluation set; monitoring quality, cost and human overrides; and handling provider model updates and retirements. It borrows most of its discipline from MLOps but rarely involves training.

### What is the difference between LLMOps and MLOps?

MLOps centers on training pipelines: data, features, retraining, a model registry and drift monitoring for models your team trained. LLMOps usually runs a model someone else trained and hosts, so behavior changes through prompts, sources and configuration, quality is measured with evaluation sets and human review, and the provider’s model updates become a change you have to manage.
