veridive is now an applied AI company. Looking for the answer engine?Looking for the answer engine? What happened

veridive TR Start a project Menu

Field notesReliability

What to monitor after an AI system goes live, and why quality slips.

Quality slips without anyone touching the system: sources change, the case mix shifts, providers update models. Watch quality, cost, speed and use together, and treat a rising override rate as the earliest warning you will get.

veridive6 min read

An AI system passes its acceptance tests, goes live and does its job. Months later nobody has touched its code, and the answers are worse. A policy was replaced, a product line was added, the provider updated the model. None of it produced an error in any log.

That is the normal life of an AI system in production, and it is why monitoring one differs from monitoring ordinary software. You watch four families of signals together (quality, cost, speed and use), each broken down by case type and language. And you treat a rising override rate, meaning people correct or reject the output more often, as the earliest warning you will get.

Why does quality change after launch?

Because the system’s surroundings move, even when the system doesn’t. The usual causes:

  • Documents are updated. A policy is rewritten or a price list replaced, and retrieval now finds the new text, or old and new versions side by side.
  • The business changes. New products, policies and channels bring case types the evaluation set never contained.
  • The case mix shifts. Month-end, campaigns and seasonal returns change the share of hard cases, so averages move even when nothing got worse.
  • The provider updates the model. A new version behind the same name, or a retirement that forces a switch, can change tone, format and accuracy.
  • Upstream systems change. A renamed ERP field, a new OCR engine or a redesigned web form changes the input the model receives.

None of these throws an exception. The system keeps answering fluently, and it is simply wrong more often for some kind of case.

Which quality signals should you track?

SignalWhat it tells youReview rhythm
Override rateShare of outputs a person rejects or substantially changes; the earliest warningTrend daily, review weekly
Edit size on draftsHow much people change before approving; slippage below the override lineWeekly
Sampled review scorePeople score a random sample against the evaluation rubric; the honest measureWeekly
Citation supportWhether cited passages support the answer; the key signal for retrieval systemsWeekly
“No source found” and escalationsA jump suggests a broken index; a sudden drop can mean guessingDaily
Validation failuresOutputs that break the schema or business rules, often after a model changeDaily
Evaluation set scoreRegression against the baselineEvery change, and monthly

Break every one of these down by case type and language. An overall average can stay flat while one product line, one channel or the Turkish-language cases get steadily worse. The breakdown is where drift becomes visible.

Which cost and speed signals matter?

SignalWhat it tells youReview rhythm
Cost per finished taskAll calls, retries and review minutes; the number the business case rests onWeekly
Tokens per task, by stepPrompts and retrieved context growing quietlyWeekly
Retries and fallbacksHidden spend, and often a quality symptomDaily
Review minutes per caseOften the largest cost line in the workflowWeekly
Slow-end latencyHow long the slowest requests take, not the averageDaily
Spend against the ceilingWhether the monthly budget will holdDaily, with alerts

Cost and quality move together more often than people expect. A model update that makes answers longer raises the bill and the review time. A retrieval change that adds context can improve answers while raising tokens per task. Read the two tables side by side.

What does usage tell you, and what doesn’t it?

Usage tells you whether the system is part of the work: the share of eligible cases that go through it, the people who use it each week, drafts opened and then discarded, and cases handled the old way instead. A team that stops using the system is giving you feedback, usually before anyone files a complaint.

What usage can’t tell you is whether the output is right. Heavy use is compatible with poor quality, especially when review has become a formality. An override rate near zero, with approvals taking seconds, deserves a check rather than a celebration.

How often should people review a sample?

Every week, by the people who own the work, using the same rubric as the evaluation set. Draw the sample at random within each case type and language, so rare but important cases always appear, and add targeted samples: case types touched by a change, new kinds of cases and outputs the system marked as uncertain. Review more heavily in the first weeks after launch and after any change, then settle into the routine.

Picture an illustrative case. A distributor’s sales-support assistant answers questions about prices and availability, citing the price list. Finance publishes a new price list for one product line as a separate file, and the old file stays in the library. From the next day, some answers for that line cite superseded prices. Nothing errors, and the overall quality score barely moves, because the line is a small share of cases. What moves is the override rate for that case type, as sales staff correct prices before replying. The weekly sample confirms it, and the cause is in the library, not the model. The fix: retire the old file, re-index, add the new prices to the evaluation set and add a check for duplicate price lists.

Findings like these should feed the routine described in building a feedback loop from overrides.

Drift doesn’t throw an error. It shows up as people correcting the system more often.

What does a one-screen dashboard contain?

Five blocks are enough:

  1. Headline numbers for each family. Sampled quality and override rate, cost per task, slow-end latency and the share of eligible cases handled, each against the threshold agreed at acceptance.
  2. The breakdown. A table by case type and language, sorted by change since the previous period, so the biggest movement sits on top.
  3. A change timeline. Every model, prompt, source and upstream change marked on the time axis, so a shift can be matched to its cause.
  4. Open alerts, each with a named owner. An alert without an owner is a notification nobody reads. Each one says who looks, how quickly and what the first step is, and serious ones follow the incident plan.
  5. Review status. How many cases were sampled, what was found and which fixes are pending.

The rhythm matters as much as the screen: a daily glance at alerts, a weekly review with the workflow owner, and a quarterly look at trends and the next improvement.

The two cheapest signals

If a system is live without any of this, start with two signals: the override rate by case type and a weekly sampled review. They cost little and catch what customers would otherwise find first. Monitoring is one of the eight conditions we check before anything goes live, and after launch, AI reliability is the service that keeps it running.

Sources

  1. Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 National Institute of Standards and Technology (NIST) nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf

Ask an assistant about this note

ReliabilityMonitoringQuality

veridive

Field notes are written and reviewed by veridive. How we write them

Questions

Questions about this note

What should you monitor in an LLM application?

Monitor four families of signals together: quality (override rate, sampled review scores, citation support, validation failures), cost (cost per finished task, tokens per task by step, retries), speed (slow-end latency) and use (share of eligible cases handled, bypasses). Break every metric down by case type and language, compare it with the thresholds agreed at acceptance, and give every alert a named owner.

Why does AI quality drift after launch?

Because the world around the system changes even when its code doesn’t. Documents and price lists are updated, new products and policies bring new kinds of cases, seasonal peaks change the case mix, providers update models, and upstream systems change the shape of the input. None of this raises an error, so drift has to be caught by monitoring and sampled human review.

What is an AI override rate?

The override rate is the share of AI outputs that a person rejects or substantially changes before using them. Tracked by case type, it is usually the earliest sign that quality has slipped, because the people doing the work notice wrong answers before any dashboard does. A rate near zero deserves a check too: it can mean review has become a formality.