From one team to many: how to roll out an AI system that worked once.
The second team is never a copy of the first: its cases, language, systems and habits differ. Re-baseline, re-run the evaluation set on its cases, find local champions and widen scope one step at a time.
veridive6 min read
The pilot team loved it. Decisions came faster, the owner signed off, and the go-live review was quick. Then the second team started, and the overrides piled up: different product lines, an older order system, and customers who write in Turkish and English in the same message.
Nothing was wrong with the system. It had simply never met this team’s work. The second team is never a copy of the first, and every difference is a place the system hasn’t been tested. So a rollout is less like copying a file and more like a short pilot per team: re-baseline, re-run the evaluation set on the new team’s cases, find local champions and widen scope one step at a time.
Why does a system that worked for one team stumble with the next?
Because its evidence came from one team. The evaluation set, the thresholds, the prompts and the review routine were all tuned to that team’s work. Four things usually differ next door:
- Case mix. Other products, other customers, other seasons, and case types the system has never seen.
- Language. Turkish, English, Arabic or a mix, with different spelling habits, abbreviations and names.
- Systems. Another ERP instance, different fields and codes, an older integration, patchier data.
- Habits. A different review culture, a manager with a different weekly rhythm, no one who was in the pilot.
The pilot also had attention the rollout won’t: builders nearby, a motivated owner and volunteers who wanted it to work. That has to be rebuilt, not assumed.
What should you check before each new team starts?
Run the same checklist for every team, and let a blank line stop the start date:
- Baseline. Volume, time per case, error rate and cost, measured for this team, not borrowed from the pilot.
- Case mix. Which case types arrive, in what proportions, and which ones are new to the system.
- Language. The languages in the inputs and outputs, and native speakers available to review.
- System differences. Instances, fields, codes and integrations, and the quality of the data behind them.
- Data access. Permissions, processing location and retention. A new legal entity or country raises questions for your data protection officer, including transfers abroad.
- Local owner. A named person who owns the workflow here and will sign off.
How do you adapt the evaluation set?
Sample the new team’s recent cases, awkward ones included, and have its own experts write the reference answers. Add them to the shared evaluation set, tagged by team and language, and run the set before go-live. Report results by team and by case type, never as one blended number.
An evaluation set only speaks for the cases it contains.
Case types that miss the bar stay with people, or run in shadow mode until they don’t. From then on, every change runs the whole set, every team’s cases included, so a fix for the new team can’t quietly break the old one.
Who needs to be involved locally?
The local owner, who defines good for this team and signs off. Local champions with protected time, who help colleagues and collect problems. The managers, because the routine that made the first team successful (a manager who asks for the new output every week) has to be rebuilt here; the note on a new everyday practice explains why. Then local IT for integrations, native-language reviewers, and the data protection contact when the team sits in another entity or country.
Training follows the same rule as the pilot: short sessions on this team’s own cases, in the language it works in. Practice like that is the core of our AI enablement work.
In what order should teams, channels or countries join?
Start where the work is closest to the pilot, then change one thing at a time: new cases, or a new channel, or a new language, never all three at once. When results drop, you will know why.
Picture an illustrative case: a returns assistant piloted with an online returns team.
- One store region. Same policy, same systems. What is new: in-store cases and photos taken by staff. Check photo quality and how the store routine fits the review step.
- The contact center. A new channel. Cases arrive as agents’ notes, less structured, and answers are needed while the customer waits. Check the case mix, response time, and whether the review step works during a live conversation.
- A second country. A different language mix, with Arabic and English and names in both scripts, local return rules and another order system instance. Check native reviewers, the local policy sources and the data questions for the DPO.
Each step gets its own cases in the evaluation set, a short shadow period, its own go-live review against the same conditions as the first, and results reported separately. The next step waits until the last one is stable.
How do you keep one system instead of many variants?
Configuration over copies. Keep one shared core: the code, the common parts of the prompts, the evaluation tooling and monitoring. Give each team local settings: language, policy sources, routing thresholds, approval limits, queue names and system connectors. Settings live in the same repository and are reviewed like code. Dashboards follow the same logic: one view, filtered by team, so owners compare like with like.
Copies look faster and cost more. A fix in one never reaches the others, the evaluation sets drift apart, and every model change has to be tested several times over. Fork only when the workflow itself is different, such as warranty claims next to returns; then it is a new workflow with its own owner and its own pilot.
Before the next team starts
Fill in the six-point checklist for your next team before announcing a date, and add its cases to the evaluation set first. Each step should pass the same go-live conditions as the first; how we work describes them. If you want a team alongside yours while the system spreads, AI reliability through a monthly Embedded AI Partnership covers monitoring, regression checks and the next improvements.
Ask an assistant about this note