How to build a call quality scorecard AI can apply consistently.
AI can review every call only against criteria it can apply the same way twice. Rewrite vague criteria such as “was empathetic” into observable behaviors, calibrate on past calls with the quality team, and cite every finding to the second.
veridive6 min read
Many contact-center scorecards have a line like “Showed empathy: 1 to 5”. Experienced reviewers score it by feel, and even they disagree: one hears warmth, another hears a script. That was tolerable when a person scored a handful of calls per agent each week. It stops working when AI scores every call, because a criterion that people can’t apply consistently will be applied inconsistently by a model too, only faster.
AI can review every call only against criteria it can apply the same way twice. So the work is in the scorecard: rewrite vague criteria into observable behaviors, decide which ones stay with people, calibrate on past calls with the quality team, and cite every finding to the exact second so it can be checked.
Why don’t human scorecards transfer directly to AI?
Because they rest on judgment the quality team built up together over years, and the form only reminds them of it. Handed to a model, the same words get the model’s own reading. Three other gaps show up quickly:
- Mixed levels. One form mixes facts (was identity verified?), judgments (was the tone right?) and outcomes decided elsewhere (was the customer satisfied?).
- Missing context. “Offered the correct resolution” needs the order record and the policy, not just the transcript.
- Transcript errors. The model scores the words it was given. Where Turkish speech recognition misheard a negation, the finding inherits the error.
Scoring every call also changes the purpose. A sample audits agents; full coverage finds patterns, such as a missing hold explanation across a whole queue.
How do you turn criteria into observable behaviors?
Apply one test to every criterion: could two reviewers, listening separately, agree on it, and point to the second where it happened or show that it didn’t? If not, rewrite it as a behavior, the condition that triggers it, and the evidence that shows it.
An illustrative rewrite. Before: “Showed empathy (1 to 5).” After, four observable statements:
- When the customer described a consequence of the problem, such as a missed workday or an extra cost, the agent acknowledged it before moving to the solution.
- The agent did not talk over the customer while they described the problem.
- When the company was at fault, the agent apologized for the specific problem, not with a generic phrase.
- Before ending the call, the agent said what happens next and when.
Write behaviors, not phrases. Acknowledgment sounds different in Turkish (“Haklısınız, bu sizin için gerçekten zahmetli olmuş”) and in English, and a phrase list misses every equivalent it doesn’t contain; give examples in both languages instead.
The result is a scorecard that fits in one table. These rows are illustrative:
| Criterion | Observable evidence | Applies to | Weight · who scores |
|---|---|---|---|
| Identity verified | Agreed details confirmed before any account information is given | Calls with account access | High · AI |
| Required disclosure | The required statement, or an approved variant, before the sale is confirmed | Sales calls | High · AI, every miss reviewed |
| Problem acknowledged | The customer’s problem restated before a solution is offered | Service calls | Medium · AI |
| Hold explained | Reason and rough length given before each hold | Calls with holds | Low · AI |
| Correct resolution | The offer matches the policy for the case type | Complaints, returns | High · AI suggests, quality team confirms |
| Vulnerable customer | Signs noticed and the agreed procedure followed | Any call where signs appear | Not scored · human only |
Which criteria should stay with people?
Anything that depends on hearing rather than reading, such as warmth, sarcasm or a customer close to tears. Anything involving vulnerable customers. Compliance determinations, where AI flags and the compliance team decides. Criteria that never reach agreement in calibration. And anything that could affect an agent’s pay or job: a finding is input to a person’s decision, never the decision.
Mark these “human only” in the scorecard, and route the calls where they apply to the quality team.
If two reviewers can’t agree on a criterion, no model will apply it consistently.
How do you calibrate on past calls?
In calibration sessions, the quality team and the AI score the same set of past calls independently, and then compare results criterion by criterion. Each difference ends one of three ways:
- People disagreed among themselves too: the criterion is unclear, so rewrite it.
- Only the AI differs: fix its instructions or add examples.
- Agreement never comes: move the criterion to human-only.
Track agreement per criterion, with agreement between people as the ceiling to aim for. Widen coverage from a sample toward every call only when agreement meets the level the quality team set, and keep scoring a regular sample by hand after launch. The illustrative call quality review on our work page follows the same path.
How should findings be shown to supervisors and agents?
Every finding links to the exact second, with the transcript excerpt and one-click playback, so anyone can check it in the time it takes to listen. That is the principle behind answers with receipts. Findings are visible to the agents themselves, not only to supervisors, and agents can dispute any finding with a reason. The quality team reviews disputes, and upheld ones become calibration examples. Findings the system is unsure about go to the quality team first, not onto an agent’s record.
For supervisors, patterns beat league tables: which behavior is slipping, in which queue, since when. Coaching works best on one or two behaviors at a time.
How do you keep the scorecard fair?
- Score only what the agent controls. Long holds during a system outage are not the agent’s fault.
- Check transcription across accents and languages. If the speech system struggles with some accents, those agents collect more false findings; compare finding rates between groups.
- Version the scorecard. Announce changes, and compare scores only within a version.
- Involve agents. People who helped write the criteria are more likely to trust the findings.
Decisions about pay, discipline or monitoring follow your HR process; questions about employee monitoring go to HR and counsel.
Rewrite the form first
Take your current form and apply the two-reviewer test to each line. Rewrite what fails, mark what stays human, and score a few dozen past calls with the quality team before any AI is involved. Calibrated review with every finding cited to the second is the core of voice and meeting intelligence, and in contact centers it sits alongside the rest of customer operations.
Ask an assistant about this note