What to monitor after an AI system goes live, and why quality slips.
Quality slips without anyone touching the system: sources change, the case mix shifts, providers update models. Watch quality, cost, speed and use together, and treat a rising override rate as the earliest warning you will get.
veridive6 min read
An AI system passes its acceptance tests, goes live and does its job. Months later nobody has touched its code, and the answers are worse. A policy was replaced, a product line was added, the provider updated the model. None of it produced an error in any log.
That is the normal life of an AI system in production, and it is why monitoring one differs from monitoring ordinary software. You watch four families of signals together (quality, cost, speed and use), each broken down by case type and language. And you treat a rising override rate, meaning people correct or reject the output more often, as the earliest warning you will get.
Why does quality change after launch?
Because the system’s surroundings move, even when the system doesn’t. The usual causes:
- Documents are updated. A policy is rewritten or a price list replaced, and retrieval now finds the new text, or old and new versions side by side.
- The business changes. New products, policies and channels bring case types the evaluation set never contained.
- The case mix shifts. Month-end, campaigns and seasonal returns change the share of hard cases, so averages move even when nothing got worse.
- The provider updates the model. A new version behind the same name, or a retirement that forces a switch, can change tone, format and accuracy.
- Upstream systems change. A renamed ERP field, a new OCR engine or a redesigned web form changes the input the model receives.
None of these throws an exception. The system keeps answering fluently, and it is simply wrong more often for some kind of case.
Which quality signals should you track?
| Signal | What it tells you | Review rhythm |
|---|---|---|
| Override rate | Share of outputs a person rejects or substantially changes; the earliest warning | Trend daily, review weekly |
| Edit size on drafts | How much people change before approving; slippage below the override line | Weekly |
| Sampled review score | People score a random sample against the evaluation rubric; the honest measure | Weekly |
| Citation support | Whether cited passages support the answer; the key signal for retrieval systems | Weekly |
| “No source found” and escalations | A jump suggests a broken index; a sudden drop can mean guessing | Daily |
| Validation failures | Outputs that break the schema or business rules, often after a model change | Daily |
| Evaluation set score | Regression against the baseline | Every change, and monthly |
Break every one of these down by case type and language. An overall average can stay flat while one product line, one channel or the Turkish-language cases get steadily worse. The breakdown is where drift becomes visible.
Which cost and speed signals matter?
| Signal | What it tells you | Review rhythm |
|---|---|---|
| Cost per finished task | All calls, retries and review minutes; the number the business case rests on | Weekly |
| Tokens per task, by step | Prompts and retrieved context growing quietly | Weekly |
| Retries and fallbacks | Hidden spend, and often a quality symptom | Daily |
| Review minutes per case | Often the largest cost line in the workflow | Weekly |
| Slow-end latency | How long the slowest requests take, not the average | Daily |
| Spend against the ceiling | Whether the monthly budget will hold | Daily, with alerts |
Cost and quality move together more often than people expect. A model update that makes answers longer raises the bill and the review time. A retrieval change that adds context can improve answers while raising tokens per task. Read the two tables side by side.
What does usage tell you, and what doesn’t it?
Usage tells you whether the system is part of the work: the share of eligible cases that go through it, the people who use it each week, drafts opened and then discarded, and cases handled the old way instead. A team that stops using the system is giving you feedback, usually before anyone files a complaint.
What usage can’t tell you is whether the output is right. Heavy use is compatible with poor quality, especially when review has become a formality. An override rate near zero, with approvals taking seconds, deserves a check rather than a celebration.
How often should people review a sample?
Every week, by the people who own the work, using the same rubric as the evaluation set. Draw the sample at random within each case type and language, so rare but important cases always appear, and add targeted samples: case types touched by a change, new kinds of cases and outputs the system marked as uncertain. Review more heavily in the first weeks after launch and after any change, then settle into the routine.
Picture an illustrative case. A distributor’s sales-support assistant answers questions about prices and availability, citing the price list. Finance publishes a new price list for one product line as a separate file, and the old file stays in the library. From the next day, some answers for that line cite superseded prices. Nothing errors, and the overall quality score barely moves, because the line is a small share of cases. What moves is the override rate for that case type, as sales staff correct prices before replying. The weekly sample confirms it, and the cause is in the library, not the model. The fix: retire the old file, re-index, add the new prices to the evaluation set and add a check for duplicate price lists.
Findings like these should feed the routine described in building a feedback loop from overrides.
Drift doesn’t throw an error. It shows up as people correcting the system more often.
What does a one-screen dashboard contain?
Five blocks are enough:
- Headline numbers for each family. Sampled quality and override rate, cost per task, slow-end latency and the share of eligible cases handled, each against the threshold agreed at acceptance.
- The breakdown. A table by case type and language, sorted by change since the previous period, so the biggest movement sits on top.
- A change timeline. Every model, prompt, source and upstream change marked on the time axis, so a shift can be matched to its cause.
- Open alerts, each with a named owner. An alert without an owner is a notification nobody reads. Each one says who looks, how quickly and what the first step is, and serious ones follow the incident plan.
- Review status. How many cases were sampled, what was found and which fixes are pending.
The rhythm matters as much as the screen: a daily glance at alerts, a weekly review with the workflow owner, and a quarterly look at trends and the next improvement.
The two cheapest signals
If a system is live without any of this, start with two signals: the override rate by case type and a weekly sampled review. They cost little and catch what customers would otherwise find first. Monitoring is one of the eight conditions we check before anything goes live, and after launch, AI reliability is the service that keeps it running.
Sources
- Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 National Institute of Standards and Technology (NIST) nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf
Ask an assistant about this note