Hours saved are not money saved: measuring the ROI of an AI system.
Time saved becomes value only when someone decides what the time is for. Measure against the baseline, attribute results carefully, and count the quality, speed and risk gains the business already puts a price on.
veridive6 min read
The quarterly review reaches the AI assistant. The slide says it saved thousands of hours. Finance asks one question: which budget line went down? Nobody can answer, and a system that probably helps starts to look like a cost.
Hours saved are not money saved. Time becomes value only when someone decides what it is for. To measure the return honestly, compare against the baseline, attribute results carefully, and count the quality, speed and risk gains the business already puts a price on.
Why do AI ROI claims fall apart under questioning?
Because they are built backwards, from usage to value. The usual weak points:
- Estimated hours. Users were asked how much time they saved; nobody measured it.
- No baseline. Nothing was recorded before launch, so there is nothing to compare with.
- Gross, not net. Review time, rework and running costs are left out.
- Everything credited to AI. A new policy, a quiet season or a new hire moved the numbers too.
- Capacity counted as cash. Freed hours are presented as savings that no budget shows.
Each is fixable, and most of the fixes happen before launch.
What should you compare against?
The baseline recorded before launch: volume, time per case, error or rework rate, cost per case and cycle time, agreed with the workflow owner. If none exists, rebuild what you can from system timestamps and logs, and measure a comparison group now.
Compare like with like: the same case types, the same season, the same team. A baseline taken in a quiet month and a result taken at month-end are not comparable, in either direction.
Done when: the owner has signed a baseline sheet that lists its sources. What goes wrong: the baseline is rebuilt from memory after launch.
How do you turn time saved into value?
Separate two things that are often merged. Capacity freed is hours now available for other work. Cost removed is a budget that actually changes: overtime, temporary staff, outsourced work or a planned hire that is no longer needed.
Time saved becomes value only when someone decides what the time is for.
Write that decision down with the owner before launch, then measure what the time went to.
| Where the time goes | What to measure | How it counts |
|---|---|---|
| Less overtime or temporary staff | The budget line | Cost removed |
| Growth absorbed without hiring | Cases per person, roles not added | Cost avoided |
| Backlog cleared, faster replies | Backlog age, response time | Service gain |
| Nothing decided | Nothing | Nothing: the hours disappear |
Done when: every freed hour has a named destination. What goes wrong: “people will focus on higher-value work”, with no measure of what that work is.
How do you count quality, speed and risk?
Count the gains the business already prices, using its own figures: rework hours, refunds and credit notes issued in error, repeat contacts, write-offs, penalties avoided, discounts captured by paying on time. Speed counts in money only where it has a price, such as a service commitment; otherwise report it as a service measure.
Risk gains, such as more consistent decisions across channels or fewer findings in audits, are real but hard to price. Report them in their own units rather than inventing a currency figure. Then subtract what the system costs to run, review and maintain; the cost of a token vs. the cost of a mistake shows how to price errors on both sides.
Done when: each gain is stated in the business’s own units, with its source. What goes wrong: a vendor’s average saving per case appears where your own measurement should be.
How do you know the system caused the change?
Three options, from simplest to strongest:
- Before and after, against the baseline. Often enough, but weak when other things changed in the same period, so list them.
- A comparison team that starts later. Rollouts are staggered anyway. Compare how the first team’s measures moved with the waiting team’s over the same weeks; the difference is the system’s effect.
- Assisted versus unassisted cases. Within one team, compare cases handled with the system’s output and without it. People reach for the tool on easy cases, so assign cases at random where you can, or compare within each case type.
Keep usage in its place. A heavily used tool can still miss its baseline: if people open the assistant for every case and then redo the draft, usage is high and time per case hasn’t moved. Usage measures adoption, not value, and belongs in monitoring alongside quality and cost.
Done when: the effect is stated with the method used and the other changes in the period listed beside it. What goes wrong: the comparison team gets the system early because it asked loudest.
What does a quarterly value review look like?
The workflow owner, a finance partner and the team that runs the system meet for an hour around one page. Take an illustrative claims team; the figures are invented, in made-up units, and are not benchmarks. At baseline it handled 3,000 claims a quarter at 12 minutes each, with 6 in every 100 reworked at 30 units each. In the reviewed quarter it handled 3,600 claims at 7 minutes each, review included, with 4 in 100 reworked. A staff hour costs 20 units; the system costs 1 unit a claim to run. At baseline speed, the extra 600 claims would have needed 120 more hours, which the owner had planned to cover with temporary staff.
| Line | Reviewed quarter | How it counts |
|---|---|---|
| Capacity freed at baseline speed | 300 hours | Capacity, not money |
| Growth absorbed without temporary staff | 120 hours, 2,400 units not spent | Cost avoided |
| Backlog of disputed claims cleared | 100 hours; oldest item from 9 days to 3 | Service gain |
| Not yet assigned | 80 hours | Nothing yet |
| Rework avoided | 72 claims, 2,160 units | Quality gain |
| Running cost | 3,600 units | Cost |
| Net financial result | 960 units |
The backlog gain and the unassigned hours are not in the net figure, and shouldn’t be. The review ends with decisions, not a score: give the 80 hours a destination, find out why time per case stopped at 7 minutes, and keep the later-starting site as the comparison group. The build cost stays in the business case, and the review checks it against the payback assumed there. What goes wrong: the meeting turns into a usage report, and nobody leaves with a decision.
Give every hour a destination
Before launch, write down what each freed hour is for and who will measure it. Quarterly improvement reviews are part of AI reliability, the baseline starts in AI strategy and discovery, and every illustrative engagement on our work page is measured against a baseline agreed at the start.
Ask an assistant about this note