Field notesStrategy & leadership
How to prioritize AI use cases when the list has thirty ideas.
Score only what you can evidence: volume, time per case, cost of errors, reachable data and a named owner. Keep value and feasibility on separate axes, because one weighted score hides the trade-off leadership actually has to make.
veridive7 min read
The workshop went well. Thirty sticky notes became a spreadsheet, the spreadsheet grew to forty rows, and every row in the priority column says “high”. There is budget, and people, for three.
A useful prioritization matrix does two things that spreadsheet doesn’t. It scores only what you can evidence (volume, time per case, cost of errors, reachable data, a named owner), and it keeps value and feasibility on separate axes, because one weighted score hides the trade-off leadership actually has to make.
Why do AI idea lists grow faster than budgets?
Ideas are free. Every department has some, every vendor meeting and hackathon adds more, and nobody is told no, because saying no feels like blocking progress. Many arrive as solutions, such as “a chatbot for HR”, rather than as work that needs changing, so they can’t be compared.
The real constraint is rarely money alone. Every idea that goes ahead needs an owner’s time to define good answers, IT’s time to open data access and reviewers’ time once it runs. Those are the scarce resources, and the matrix exists to spend them on the few ideas most likely to pay back.
Which ideas should be filtered out before scoring?
First, rewrite each idea as a workflow: who does what, for which cases. “AI for customer service” becomes “draft replies to delivery-status emails for the retail support team”.
Then run the four tests from choose the workflow before the model: is it frequent, is it bounded, is the data reachable, does someone own it? An idea that fails one waits on a parked list with the reason beside it, such as “no owner” or “data in personal folders”, and comes back when the reason is fixed.
Also remove ideas that don’t need AI at all. An exact calculation, a lookup or a broken form is a job for code or a redesign, and scoring it would only flatter it.
What should each use case be scored on?
Six criteria, three for value and three for feasibility, each scored from 1 to 5. Describe the anchors in words, so two people scoring the same idea land close together. On feasibility, a high score means easier or safer.
| Criterion | Scores 1 | Scores 3 | Scores 5 |
|---|---|---|---|
| Value: volume | A few cases a month | Dozens a week | Hundreds a day |
| Value: time per case | Under a minute | Several minutes | An hour or more |
| Value: cost of errors now | Rare and cheap to fix | Rework and repeat contacts | Losses, complaints or compliance exposure |
| Feasibility: data access | In heads or personal folders | In approved systems, access not granted | In approved systems, read access agreed |
| Feasibility: integration | Writes into several core systems | Reads one or two systems, writes drafts | Stands alone or reads one system |
| Feasibility: risk | Wrong output reaches customers or money | Errors caught by review first | Internal, reversible and reviewed |
Cost of errors counts what mistakes cost in the current process, the prize for reducing them. Risk counts what a wrong output from the new system could do. Add each group separately, so both axes run from 3 to 15. If leadership wants weights, apply them within an axis, never across it. Deciding that cost of errors matters more than volume is a fair judgment; blending feasibility into value is how the trade-off disappears.
Three house rules keep the scores honest:
- No owner, no score. An idea nobody will run is not a use case yet.
- Score with the people who do the work, not only with the idea’s sponsor.
- Write the evidence next to every score, with its source, so anyone can challenge it.
How do you score value without inventing numbers?
Use counts, not impressions. Volume comes from the ticketing system, the ERP or the shared mailbox. Time per case comes from a short time study: sit with the team for a morning and time twenty real cases. The cost of errors comes from rework logs, credit notes and complaints.
When there is no evidence, don’t guess a 3. Score it 1 and add it to the list of things discovery must measure; an unknown scored as average quietly pushes weak ideas up.
Leave money out at this stage. Converting hours into currency starts arguments about rates and hides whether the evidence was any good. Money belongs in the business case for the shortlisted few.
How do you score feasibility and risk?
Ask the people who will have to deliver it. IT scores data access and integration with the owner of each source system; security and the data protection officer weigh in on risk, including whether personal or special-category data is involved.
Score risk as the design would run, not as a worst case: an output a person approves before anything happens is lower risk than the same output sent straight to a customer.
One weighted score hides the trade-off leadership actually has to make.
Then keep the axes apart. An idea with high value and low feasibility and one with low value and high feasibility can share a total, yet they need opposite decisions: fix a blocker, or start now.
What does a filled-in matrix look like?
Take an illustrative distributor with five workflows that passed the filter. The scores are invented for the example.
| Use case | Value | Feasibility |
|---|---|---|
| A. Supplier confirmations checked against open orders | 12 | 12 |
| B. Replies to delivery-status emails drafted for agents | 11 | 10 |
| C. Sales contracts checked against standard clauses | 13 | 5 |
| D. Month-end commentary drafted for the controller | 6 | 12 |
| E. Employee policy questions answered with sources | 7 | 10 |
Behind each number sits a line of evidence. For C’s value score: dozens of contracts a week (from the CRM), an hour or more each (timed with two account managers), and a clause dispute that ended in a credit note (from the finance log).
Plot value up and feasibility across, and each corner has a rule:
- High on both: start here. That is A.
- High value, low feasibility: fix the blocker first. C’s contracts sit in personal drives, so its first project is one approved repository.
- Low value, high feasibility: a general assistant or a better report will do. That is D.
- Low on both: drop it.
The shortlist: A now, B next once its review step for customer replies is designed, and C’s data work alongside. E waits for capacity. Added into one total, C and D tie at 18, though one needs months of groundwork and the other could start at once.
The matrix is also a CSV sheet, a starting point you can adapt: the four filter tests, the six criteria with an evidence column for each, both totals, the decision and three of the rows above.
When should the scores change?
Whenever an estimate becomes a measurement. A Discovery Sprint replaces guessed volumes and times with a baseline; a pilot replaces guessed feasibility with evidence about data, integration and review time. Rescore then, and expect some ideas to fall. Keep the old scores beside the new ones: an idea that dropped from 4 to 2 once measured teaches the next workshop more than any template.
Scores also move when a blocker goes, such as an owner named or access granted, and when the business changes, with a new product line or channel. Review the whole list at a fixed rhythm, quarterly is usually enough, and let it feed the now, next and later of your AI roadmap.
From sticky notes to workflows
Rewrite your list as workflows and apply the four tests in one sitting, then score the survivors with the people who do the work. The opportunity map from AI strategy and discovery is this matrix built on measured evidence, and an Executive Build Day ends with three prioritized opportunities.
Ask an assistant about this note