veridive is now an applied AI company. Looking for the answer engine?Looking for the answer engine? What happened

veridive TR Start a project Menu

Field notesCustomer & commerce

How to design an intent taxonomy that AI routing can live with.

Routing fails more often because of the categories than because of the model. Fewer, mutually exclusive intents, each with examples and an owner, make AI routing measurable.

veridive5 min read

The routing pilot looked promising until someone checked the labels. The model had sent a “where is my refund?” email to the returns queue, and half the team said it belonged with payments. Nobody could prove either side, because the old list had both “Refund status” and “Return – refund query”, and agents had used them interchangeably for years.

That is the usual shape of the problem. Routing fails more often because of the categories than because of the model. Fewer, mutually exclusive intents, each with examples and an owner, make AI routing measurable, and the template below can be filled in by a team lead in an afternoon.

Why do routing projects stall on categories?

  • Overlap. When two categories both fit, the model’s “error” is a coin toss between them, and no amount of tuning fixes it.
  • Too many codes. A hundred reason codes, of which agents use a handful consistently and the rest occasionally, by habit.
  • Mixed purposes. One list tries to serve routing, reporting and root-cause analysis at once, so a single code carries three facts.
  • No owner. A category with no queue behind it means a correctly routed request lands nowhere.

A model evaluated against labels people don’t agree on can’t be measured at all.

How many intents should you have?

As few as still change what happens next. An intent earns its place when it leads to a different queue, a different handling process or a different urgency. If two intents go to the same team and are handled the same way, merge them for routing and keep the distinction elsewhere. Check volume too: an intent that receives only a trickle of requests can’t be measured, so fold it into a broader intent or into “other” until it grows.

That “elsewhere” is a second level: a detail field for reporting, which can be longer and can be suggested by AI when the case closes. Routing needs a list short enough for a new agent to learn; reporting can have as much detail as the analysts will actually use.

How do you write an intent definition?

One row per intent, in a shared document the whole team can read:

FieldWhat to writeExample
NameShort, from the customer’s point of viewRefund status
DescriptionOne sentence on what the customer wantsAsks when or whether a refund will arrive for an approved return or cancellation
Include examplesReal messages, including awkward ones“Where is my refund?”, “İadem onaylandı ama para hesabıma geçmedi” (return approved, no money yet)
Exclude examplesNear misses, and where they go instead“I want to send this back” goes to Return request
Owner queueThe team that handles it, and a named ownerPayments desk
UrgencyThe default, and what raises itRaised if the customer mentions a chargeback
LanguagesLanguages it arrives in, with examples in eachTurkish, English

The exclude examples do the most work. They are where overlaps get settled, once, in writing.

What happens to requests that fit nowhere?

They go to “other”, which routes to experienced agents. The bucket is necessary, and it is also where the model should send cases it can’t place with confidence.

What keeps it from becoming a dumping ground is a weekly review. A named person reads the week’s “other” requests and sorts them: a recurring pattern becomes a candidate intent, a request that should have matched points to a definition that needs work, and the rest stay put. If the bucket keeps growing, the taxonomy is due for a review.

Messages with two requests, such as “where is my order, and can I change the address?”, need a rule too: route by the one that needs action first, and record the other.

How do you test the taxonomy before any model sees it?

Give two people the written definitions and the same sample of real requests, a few hundred across channels, and have them label it separately. Then compare.

A model can’t be more consistent than two people reading the same definitions.

Every disagreement is a finding. Usually it means two intents overlap, a description is vague or an exclude example is missing. Fix the definitions and label a fresh sample until agreement is high and stable. The agreement you reach is the realistic ceiling for the model, and the labeled sample becomes the start of your evaluation set, which our evaluation set template helps structure.

How do you change it without breaking reports?

Version the taxonomy, keep a change log of what changed and why, and keep a mapping from every old reason code to a new intent. Re-run the evaluation set, update routing rules and adjust dashboards in the same release. Every service metric reported by request type depends on the taxonomy, as our note on customer service AI metrics shows.

In an illustrative case, an e-commerce support team with a long list of legacy reason codes cuts it to a short list of intents: order status, delivery problem, return request, refund status, payment issue, product question, account and other. Most legacy codes map to exactly one new intent, so “Kargo gecikmesi” (shipping delay) and “Late delivery complaint” both become delivery problem; the few that mixed returns and refunds are split by a written rule. Trend reports read the old codes through the mapping, so the returns line continues across the change instead of starting again from zero.

Label a sample twice

Export a sample of recent requests with the codes agents chose, and ask two people to label it against a draft list. Intent routing is one of the first steps in customer operations, and the routing itself is usually built as custom AI software around your ticketing system and queues.

Ask an assistant about this note

Customer & commerceRoutingCustomer service

veridive

Field notes are written and reviewed by veridive. How we write them

Questions

Questions about this note

What is an intent taxonomy in customer service?

An intent taxonomy is the list of reasons customers contact you, each defined clearly enough to route a request to the right team. A good one has a short list of mutually exclusive intents, each with a description, include and exclude examples, an owner queue, an urgency rule and the languages it arrives in, plus a reviewed “other” bucket.

How many intents should a support team have?

As few as still change what happens next. An intent earns its place when it leads to a different queue, a different handling process or a different urgency. Distinctions that matter only for reporting belong in a separate detail field. The list should be short enough for a new agent to learn and for two people to apply consistently.