The gap between 66% and 20%

Deloitte's State of AI research found that 66% of organisations report efficiency gains from AI, and 20% report increased revenue.

That gap is the whole problem, and it isn't a technology problem. Those organisations got what they paid for. Systems ran faster, work took less time, throughput improved. The efficiency was real.

It just never became money.

PwC's January 2026 survey of CEOs found the same pattern from the other end: 56% reported neither revenue growth nor cost reduction from AI over the previous twelve months, and only 12% achieved both. Separately, fewer than a third of organisations say they can measure AI ROI with confidence at all, and only around 39% attribute any financial impact to it.

The standard interpretation is that AI is overhyped. The more useful interpretation is that efficiency and revenue are connected by a mechanism most companies never build.

Efficiency creates capacity. Capacity is not money.

Here is the step that gets skipped.

An automation saves your operations team twenty hours a week. That is a real, measurable, defensible efficiency gain. It is also worth exactly nothing on your P&L until one of three things happens:

  1. The capacity is reallocated to work that generates revenue — and someone decides what that work is.
  2. The headcount changes, through reduction or, more commonly, through not hiring the next person you were about to.
  3. The throughput increases and there is genuine demand to absorb it — you now process 400 orders a day instead of 200, and 400 orders exist.

If none of those happens, the twenty hours dissipate. They become slightly less pressure, slightly more slack, meetings that expand to fill the time. Everyone reports the project as a success, because by the metric it was measured on, it was.

Technology creates the conditions for improved performance, not improved performance itself.

This is why the AI value gap persists even in organisations that deployed competently. The conversion is a management decision, not a technical outcome, and it almost never happens automatically.

The practical consequence for anyone planning a technology initiative: decide the conversion mechanism before you build, not after. Which of the three routes applies? Who owns making it happen? What will you do with the capacity? A project with no answer to that question is buying efficiency and hoping.

Cost per seat is not an ROI number

The measurement layer has been shifting, and the direction is clear. The Futurum Group's 1H 2026 survey of 830 IT decision-makers found "productivity gains" declining as the top AI ROI metric while direct financial impact nearly doubled to 21.7%.

For the first wave of AI investment, productivity metrics were the right currency. The technology was new, adoption was the challenge, and proving people used it and went faster was sufficient. That phase is over. Boards and CFOs are now asking for cost per outcome, not licence counts.

The distinction matters more than it sounds:

  • Cost per seat tells you what you spent. It's an input.
  • Cost per outcome tells you what you got. Cost per resolved ticket, per processed invoice, per qualified opportunity, per completed workflow. It's a unit economic, and it can be compared against the manual alternative.

Getting to cost per outcome requires something most teams don't have: a counterfactual. What would this same output have cost without the system? Answering that requires a baseline measured at deployment, not reconstructed afterwards from memory and optimism.

That single practice — measuring the before state before you build — is the difference between being able to prove return and being able to estimate it. It costs almost nothing at the start of a project and is close to impossible to recover later.

What to baseline, concretely: current cost per unit of work, current cycle time end to end, current error and rework rate, current volume, and current headcount allocation to the process. Five numbers. An afternoon's work. Skipped by most teams, and their absence is why so many organisations can describe their AI programme fluently and cannot value it.

The attribution trap

There's an opposite failure worth naming, because it's the one that destroys credibility rather than value.

One of the most commonly cited obstacles to proving AI value is the difficulty of separating AI impact from overall business growth. Revenue went up. AI was deployed. The temptation to draw a line between them is considerable, especially when a budget renewal depends on it.

Don't. Overclaimed attribution survives exactly one CFO who checks. Once the finance team has caught one inflated number, every subsequent claim from technology gets discounted — including the true ones.

The defensible approach is narrower and holds up better:

  • Attribute at the unit level, not the aggregate. "Cost per invoice processed fell from £4.10 to £0.90" is checkable. "AI drove 12% revenue growth" is not.
  • State the counterfactual explicitly. What you compared against, over what period, and what else changed.
  • Separate the three value types — cost reduction, capacity gain, revenue enablement — and don't add them together as though they were the same currency. Cost reduction hits the P&L directly. Capacity gain only converts if reallocated. Revenue enablement is slower and harder to isolate.
  • Report what you can't isolate as what it is. "We can't separate this from seasonality" is a more credible sentence than a confident number that turns out to be wrong.

Worth applying the same scepticism to the industry's own statistics. The widely repeated "95% of GenAI pilots fail" line comes from a study measuring a specific kind of deployment against a six-month P&L test — a window short enough to make most enterprise software look like failure. Numbers this convenient usually measure something narrower than the headline suggests, and repeating them uncritically in a board paper is its own credibility risk.

Three routes from capability to revenue

Technical capability converts to money in one of three ways. Each has a different timeline, a different metric, and a different owner. Confusing them is why so many business cases collapse under scrutiny.

Route 1 — Lower cost to serve

The most direct and the easiest to prove. The same output for less money: fewer hours per unit, less rework, less escalation, lower error cost.

Metric: cost per unit of work, before and after.

Timeline: fastest — often visible within a quarter.

Trap: it only reaches the P&L if the freed cost is actually removed or redeployed. Otherwise it's capacity, not saving.

Route 2 — Increase throughput at fixed cost

Handle more volume without proportionally more people. This is the route most automation projects are actually on, whether or not anyone says so.

Metric: volume per period at constant headcount; revenue per employee.

Timeline: medium — needs demand to absorb the capacity.

Trap: useless without demand. Doubling processing capacity in a business constrained by lead flow improves nothing, and you've automated the wrong end of the pipe. Check where the constraint actually sits before you invest.

Route 3 — Enable something you couldn't sell before

New capability, new segment, new pricing power, faster response times that win deals you previously lost.

Metric: win rate, deal size, expansion revenue, retention.

Timeline: slowest, and the hardest to attribute.

Trap: the most exciting route and the one most likely to be assumed rather than evidenced. If nobody has said they'd pay for it, this is a hypothesis.

The sequencing that works: pursue near-term efficiency wins to fund longer-term transformation, rather than expecting all three simultaneously. BCG's research points the same way — leaders reinvest early returns into stronger capability, which compounds, while laggards wait for the transformational payoff that was never going to arrive first.

The metrics worth putting on the board slide

Leading indicators (first 90 days): adoption and usage against expectation, cycle time on the target process, error and rework rate, escalation rate, cost per unit of work.

Lagging indicators (two to four quarters): revenue per employee, gross margin on the affected service line, cost per outcome versus the baseline, net revenue retention if the capability touches the customer, and win rate if it touches sales.

The one most teams miss: the cost of the system itself, fully loaded. Model usage, infrastructure, maintenance and the engineering time to keep it running. A project reporting savings while omitting its own run cost is not reporting a return.

Who owns it

The structural finding across this research is consistent and unflattering: most enterprises operate without clear executive ownership of AI returns, and the result is fragmented investment, value leaking through poor adoption, and returns that cannot be validated even when they exist.

Boards have stopped asking whether the technology works. They're asking who owns the return — and in most organisations the honest answer is nobody. Technology owns delivery. Finance owns measurement. Operations owns the process. The conversion from capability to money falls between them.

Large organisations are formalising this, with some standing up value realisation functions whose job is to tie investment to strategy and force one question: did the business case pay off? For a smaller company the apparatus is unnecessary and the discipline isn't. It reduces to naming, for each initiative, one person accountable for the business outcome — not the delivery — and asking them the same question at a fixed interval.

When revenue is the wrong measure

A caveat, because forcing revenue attribution onto everything produces fiction.

Some technical work should be justified on other grounds: risk reduction, compliance obligation, security posture, table stakes to remain credible in a category, or removing a constraint that blocks something else. These are legitimate investments with no clean revenue line, and demanding one produces invented numbers that damage the credibility of every real number around them.

The honest framing is to say which category an initiative belongs to at the point of approval. This one pays back in cost per unit. This one is a compliance requirement. This one removes a constraint and its value depends on what we do next. Three defensible cases beat one uniform fiction.

What to do this week

  1. Take your last completed technical initiative and ask which of the three routes it was on. If the answer isn't obvious, that's the finding.
  2. For your next one, write down the five baseline numbers before anything gets built — cost per unit, cycle time, error rate, volume, headcount allocation.
  3. Name the conversion mechanism. If this succeeds and creates capacity, what specifically happens to that capacity, and who decides?
  4. Name one accountable person per initiative — for the outcome, not the delivery.
  5. Audit your current reporting for cost per seat masquerading as ROI. Replace it with cost per outcome, or state plainly that you can't yet.
  6. Check the constraint. If the process you're about to automate isn't the bottleneck, the efficiency gain will be real and the business impact won't.

The uncomfortable summary: most technology investments that fail to show returns didn't fail technically. They produced exactly what was asked of them, and nobody had decided in advance what that would be worth or who would make it so. That decision is cheap to make at the start and nearly impossible to reconstruct at the end.

Sources: Deloitte, "State of AI in the Enterprise" (2026, 3,235 leaders across 24 countries); PwC, 2026 Global CEO Survey (4,454 CEOs across 95 countries); Futurum Group, 1H 2026 IT decision-maker survey; BCG research on AI reinvestment.

Want to baseline your next initiative before you build it, so you know whether it pays back? Get in touch.