# How to measure AI in customer service beyond handle time.

> In this field note, veridive explains how to measure AI in customer service beyond average handle time. It offers a metrics template covering resolution, repeat contacts, cited-answer accuracy, draft edits, escalations and satisfaction by request type and channel, shows how to count deflection honestly, and sketches a weekly dashboard compared with the pre-launch baseline.

Average handle time can fall while customers call back. Measure resolution, repeat contacts, the accuracy of cited answers, how much agents edit drafts, escalations and satisfaction by request type, against the baseline recorded before launch.

## Key takeaways

- Handle time measures how fast a contact closed, not whether the problem went away; read it only next to outcome metrics.
- Track resolution, repeat contacts, answer accuracy, draft edits, escalations and satisfaction by request type and channel, against the pre-launch baseline.
- Count deflection only when the customer didn’t come back about the same issue through any other channel.
- Agent edits and override reasons are early quality signals, but a falling edit rate needs a quality sample to confirm it.

A new assistant goes live in the contact center, and the first weekly report looks excellent: average handle time is down. A few weeks later, the returns team notices more customers writing again about orders they had already asked about. Both observations are true. Only one of them was on the dashboard.

Handle time measures how fast a contact closed, not whether the customer’s problem went away. To measure AI in customer service honestly, track resolution, repeat contacts, the accuracy of cited answers, how much agents edit drafts, escalations and satisfaction, split by request type and channel, and compare each with the baseline recorded before launch. The template below can be rebuilt in an afternoon.

## Why can handle time mislead?

Three mechanisms move it for reasons unrelated to quality:

- **Fast but unfinished.** A draft that answers half the question closes quickly and produces a second contact. Both contacts are short, so the average improves while total work grows.
- **A changing mix.** When self-service takes the easy requests, agents keep the hard ones, and their average rises even if the assistant helps them.
- **Targets change behavior.** Closing early and transferring are both quick ways to hit a handle-time number.

Handle time is still worth tracking as an efficiency signal. It just can’t be the headline.

> Handle time tells you how fast a contact closed, not whether the problem went away.

## Which outcome metrics matter most?

Write one definition per metric before launch, and use the same definition for the baseline.

| Metric | Definition | What it catches | How it can be gamed |
|---|---|---|---|
| **Resolution** | Solved with no further contact on the same issue within an agreed window | Answers that sounded right but weren’t | Closing early; a window too short to see the callback |
| **Repeat contacts** | Another contact from the same customer about the same issue, in any channel | Problems that closed but didn’t go away | Counting only within the original channel |
| **Cited-answer accuracy** | Sampled answers whose facts match the cited source and current policy | Wrong or outdated information | Sampling only easy request types |
| **Draft edit rate** | Drafts agents change materially before sending | Drafts that look plausible but need work | Agents sending without reading |
| **Escalations** | Cases handed to a person or a higher tier, with the reason | Requests the AI can’t handle well | Hiding the route to a person |
| **Satisfaction** | Post-contact rating by request type | How the answer felt | Asking only after resolved contacts |
| **Handle time** | Agent time per contact | Efficiency | Early closes and transfers |

Report every metric by request type and by channel: email, chat, WhatsApp and phone. An overall average hides exactly the place where the AI fails, such as returns questions on WhatsApp. Stable request types need a stable list of intents, which our note on [intent taxonomies](https://veridive.com/insights/intent-taxonomy-for-ai-routing/) covers.

## How do you measure answer quality?

Review a sample every week, per request type, against the source each answer cites. Reviewers check four things: the fact is correct, the cited source is the right and current one, the policy is applied correctly, and the answer is complete. A short rubric keeps two reviewers scoring the same way. Oversample the request types where a wrong answer is expensive, such as refunds and complaints.

Include cases where the right answer is “I don’t know”. A system that answers confidently when the knowledge base is silent should score worse than one that hands over.

For calls, the same rubric can run on transcripts, with each finding linked to the moment in the recording, as in [voice and meeting intelligence](https://veridive.com/solutions/voice-meeting-intelligence/). Our note on [call quality scorecards](https://veridive.com/insights/call-quality-scorecard/) shows how to write the criteria.

## What do agent edits tell you?

Every draft an agent sends is a vote. Record it in four buckets: sent as drafted, lightly edited, heavily edited or rejected. Add a short pick-list of override reasons: wrong fact, outdated policy, missing information, wrong tone, wrong request type, context the draft didn’t have.

The reasons point to the fix. Outdated policy is a job for the content owner. Missing information often means retrieval didn’t find the right article. A wrong request type is a routing problem, not a drafting one.

Treat a falling edit rate with care. It can mean the drafts improved, or that agents stopped reading them. The weekly quality sample tells the two apart: if reviewers find errors that agents sent unchanged, review has stopped being real.

## How should deflection be counted honestly?

Deflection is the share of requests resolved through self-service without an agent. It counts only when the customer didn’t come back about the same issue through another channel within the agreed window. A chat that ends politely and is followed by a phone call is not a deflection. It is a failure that also cost the customer time.

- **Link contacts across channels** by phone number, email address or order number, and accept that matching will be imperfect.
- **Count abandonment separately.** A customer who closes the chat without an answer was not deflected.
- **Show containment and resolution side by side.** Containment says the customer never reached an agent; resolution says the problem went away.

## What does a weekly dashboard look like?

One screen, one row per major request type, a channel filter, and every figure next to its baseline: volume, resolution, repeat contacts, edit rate, escalations, satisfaction and handle time. Below it sit the quality sample findings and the top override reasons.

Here is an illustrative example, in invented numbers. After launch, average handle time for returns questions falls. The same row shows repeat contacts rising, and the override reasons explain why:

| Returns questions (invented numbers) | Baseline | After launch |
|---|---|---|
| Average handle time, minutes | 9 | 6 |
| Repeat contacts within seven days, per 100 | 12 | 19 |
| Drafts heavily edited or rejected, per 100 | – | 31 |
| Top override reason | – | Refund timing missing |

The drafts explained how to send the item back but not when the refund would arrive. Some agents added the timing by hand and flagged it; where they didn’t, customers asked again. The fix was a content fix, not a model change: the refund-timing article was updated and made a required source for returns drafts. Handle time alone would have called the launch a success.

## Definitions and a baseline

Before launch, write the definitions, choose the window for repeat contacts and record the baseline by request type and channel. The [customer operations](https://veridive.com/solutions/customer-operations/) page shows where these measures fit in a first project, and keeping them honest after launch is part of [AI reliability](https://veridive.com/services/ai-reliability/) work.

## Frequently asked questions

### How do you measure the success of AI in customer service?

Measure outcomes, not only speed: whether requests are resolved without a repeat contact, how accurate the cited answers are, how much agents edit AI drafts, how often cases escalate and how satisfied customers are, split by request type and channel. Compare each figure with the baseline recorded before launch, and read handle time only alongside them.

### What is a deflection rate, and how should it be counted?

A deflection rate is the share of requests resolved through self-service without an agent. It should count a request only when the customer didn’t come back about the same issue through another channel, such as phone or email, within an agreed window. Abandoned chats and conversations followed by a call are failures, not deflections.

### Why can average handle time mislead after an AI launch?

Because a contact can close faster without solving the problem. Incomplete answers produce repeat contacts that each look short, so the average falls while total work rises. Automation can also move easy requests away from agents, which raises their average for reasons unrelated to quality. Handle time is useful only next to resolution and repeat-contact figures.
