What a support agreement for an AI system should promise.
You can’t promise accuracy the way you promise uptime. A useful agreement commits to quality thresholds on the evaluation set, incident response times, rules for model and prompt changes, cost reporting and a regular improvement review.
veridive5 min read
The draft support agreement promises what IT agreements always promise: availability, and response times for tickets. An AI system can meet both and still be wrong for a month, because most of its failures are not outages. They are answers that look right.
A useful AI support SLA commits to what can be measured instead: quality thresholds on an agreed evaluation set, response times for incidents, rules for changing models and prompts, cost reporting and a regular improvement review. The clause checklist is at the end.
Why do standard IT service levels fit AI systems badly?
They measure availability, and availability is rarely the problem. Quality depends on inputs the supplier doesn’t control, such as your documents and your case mix, and on a model provider neither party controls. The system can change without a release when that provider updates a model. And a single accuracy percentage is either meaningless or unenforceable: accurate on which cases, measured by whom, over what period?
Keep the availability terms. Add clauses for the way AI systems actually fail.
What can be promised about quality?
You can’t promise accuracy the way you promise uptime.
What you can promise is a method. Tie the commitment to an evaluation set both sides have agreed, versioned and kept current, and to a sampled review of live cases, because no set anticipates everything production brings.
Take an illustrative clause for a returns assistant: Before any release, the system meets the agreed thresholds on the current evaluation set for every case type, and never proposes a refund above the order value. Each week, the supplier and the returns lead score a random sample of forty live cases against the agreed rubric. If two weekly samples in a row fall below threshold, the supplier investigates, reports the cause and proposes a fix within an agreed time, and either side may roll back to the last good version.
Compare that with a single accuracy percentage. The clause says what is measured, on which cases, by whom, and what happens next.
What should incident response times cover?
Define severities by impact rather than by component, as in our note on an incident plan for the first hours, and agree response expectations per level that match the workflow’s hours: a customer-facing assistant that runs at night needs different cover from a back-office queue. Make sure the triggers include AI-specific failures, such as harmful or off-policy output, data shown to the wrong person, quality below threshold and cost spikes, not only downtime.
Commit to times to acknowledge and to contain, not to fix, since a fix has to pass the evaluation set first. State that your team may stop the system without waiting for the supplier, and that anything involving personal data reaches your DPO at once.
How should changes to models and prompts be handled?
Write down who may change what, and how each change is approved:
- The model provider updates or retires models, and neither party controls when. The supplier watches for this, pins versions where possible and re-runs the evaluation set before adopting a new one.
- The supplier changes prompts, retrieval settings and model choices. Each change passes the evaluation set and gets a change note, and significant ones need your approval.
- Your team changes source documents, policies and business rules, which can change answers just as surely. Agree who tells whom.
Keep a change log both sides can read, and a tested way to roll back.
What about cost reporting and ceilings?
Ask for a monthly report of cost per task, by feature, with volumes and anything unusual explained. Agree the ceiling, what happens as spend approaches it, and who pays if it is exceeded; that last point goes to counsel. If several workflows share one provider account, agree how costs are split between them. Ask for notice before provider price changes are passed through. Where possible, hold the model provider account in your own name, so costs stay visible and the account stays with you if the supplier changes.
What should the regular review include?
A monthly or quarterly review with the workflow owner: quality trends by case type, incidents and the evaluation cases they produced, changes made, cost per task, overrides and feedback, and the next improvements, chosen on evidence. Agree in advance what the review may decide, such as widening the scope, adding a language or retiring a feature, and who signs it off. The review is where the evaluation set grows; our note on what to monitor after launch lists the signals worth bringing.
The clause checklist
- Scope: the system, workflows, languages, hours and integrations covered, so gaps are visible.
- Quality measurement method: the evaluation set, thresholds per case type and the sampled review, because quality is what usually fails.
- Incident severities: triggers, response expectations and who may stop the system, because AI failures are rarely outages.
- Change management: who may change models, prompts and sources, how changes are tested and approved, and how to roll back, because the system can change without a release.
- Cost reporting: cost per task, the ceiling and what happens near it, so every increase has an explanation.
- Review rhythm: how often, with whom and what it produces, because that is where the system improves.
- Exit and handover: what you receive, from code and prompts to evaluation sets and runbooks, so you can change supplier; our handover checklist lists it.
Contract terms go to counsel; these are the questions to bring. Agreements along these lines are part of our AI reliability work, usually through an Embedded AI Partnership with service levels agreed together.
This note is general information, not legal advice.
Ask an assistant about this note