The ROI credibility crisis
Here's a scene that plays out in boardrooms every week:
Someone presents an AI initiative. The ROI slide shows a 10x productivity improvement. The CFO raises an eyebrow. The slide gets politely acknowledged and immediately discounted.
Why? Because everyone in that room has seen the same slide before. For RPA. For cloud migration. For the CRM rollout three years ago. And not once did the actual results match the projection.
The AI industry has an ROI credibility problem. Not because the returns aren't real — they often are. But because the way we measure and present them has been contaminated by years of overclaiming.
If you want your AI investment case to survive contact with a skeptical finance team, you need a framework that's honest, conservative, and grounded in the metrics they already care about.
The four honest ROI categories
Forget "productivity multipliers" and "efficiency gains." These are handwaving. Here are four categories of return that can be measured, verified, and defended:
1. Cost avoidance
This is the simplest and most credible category. It answers the question: "What spending are we avoiding by having this capability?"
Examples:
- An AI system generates compliance reports that would otherwise require a consultant at $200/hour. Measure: hours of consulting avoided × rate.
- An AI system monitors project data that would otherwise require a dedicated analyst. Measure: salary equivalent of the monitoring role.
- An AI system automates data reconciliation that currently requires three days of manual work per month. Measure: time × loaded cost of the people doing it.
Cost avoidance is credible because it maps directly to existing line items. The CFO can see the invoice for the consultant you no longer need.
2. Time recovery
Time recovery is different from "productivity improvement" — because it's specific and measurable. It answers: "How many hours per week are being returned to people who were previously spending them on low-value tasks?"
The key difference: you're not claiming people work 10x faster. You're claiming a specific, measurable amount of time was being spent on a specific task, and that time is now recovered.
Examples:
- Weekly client update preparation went from 3 hours to 20 minutes per project manager. Time recovered: 2 hours 40 minutes per PM per week.
- Invoice reconciliation went from 2 full days to 45 minutes per month. Time recovered: 14.25 hours per month.
- Audit preparation went from 2 weeks to 3 days. Time recovered: 7 days per audit cycle.
Measure by timing the process before and after. Not sampling. Not estimating. Actual time tracking on actual tasks.
3. Quality uplift
This category is harder to quantify but often the most valuable. It answers: "What errors, inconsistencies, or omissions are we preventing?"
Examples:
- Compliance reports previously had a 12% error rate requiring rework. AI-generated reports have a 2% error rate. Quality uplift: 10 percentage points of error reduction.
- Client-facing documents had inconsistent formatting across 8 templates. AI-generated documents use enforced templates. Consistency: 100%.
- Data entry had a 5% error rate. AI-assisted data entry (with human verification) has a 0.5% error rate.
Quality uplift is credible when you can show the before-and-after metrics. It's especially powerful in regulated industries where error rates have direct cost and compliance implications.
4. Capability unlock
Some AI returns aren't about doing existing things better. They're about doing things that weren't possible before.
Examples:
- Cross-referencing data from five systems in real-time to identify at-risk projects. This analysis was theoretically possible before, but practically impossible — nobody had the time to pull data from five systems, normalise it, and run the comparison weekly.
- Generating customised client summaries at the end of every sprint. This never happened before because the effort-to-value ratio was too high.
- Running scenario analysis across the entire project portfolio. This was a quarterly exercise at best. Now it's on-demand.
Capability unlock is the hardest to put a dollar value on, but it's often the ROI category that resonates most with executives. It's not "we saved X" — it's "we can now do Y, and we couldn't before."
The measurement framework
For each AI initiative, build a simple measurement table:
| Category | Metric | Before | After | Delta | Source |
|---|---|---|---|---|---|
| Cost avoidance | Monthly consulting spend | $8k | $2k | $6k saved | Invoice records |
| Time recovery | Hours on reporting per week | 15 | 3 | 12 hrs recovered | Time tracking |
| Quality uplift | Report error rate | 12% | 2% | 10pp improvement | QA review logs |
| Capability unlock | Cross-system risk analysis | Not possible | Weekly | New capability | N/A |
The "Source" column is what makes this credible. Every number is traceable to an existing system of record. No models. No projections. No "we estimate."
The presentation rules
When presenting AI ROI to leadership:
1. Lead with cost avoidance. It's the most immediately credible.
2. Show time recovery in hours, not percentages. "12 hours per week recovered" is more believable than "80% faster."
3. Use quality uplift in regulated contexts. If your audience cares about compliance, error reduction is gold.
4. Save capability unlock for the strategic conversation. It's the most exciting, but save it for after credibility is established.
5. Never combine categories into a single number. A blended "total ROI" number loses credibility because it mixes solid metrics with softer ones.
The bottom line
AI delivers real returns. But the moment you inflate them, you lose the audience. Measure what's measurable. Cite your sources. Present conservatively. And let the actual results build the case for expansion.
The best AI ROI story isn't the one with the biggest number. It's the one the CFO believes.