The most-quoted number in AI product strategy is being read wrong
You have seen the statistic. Ninety-five per cent of enterprise AI pilots fail. It has been cited by Fortune, Forbes and Harvard Business Review, and it is still being used in 2026 as the opening line of pitch decks — often by companies selling the very thing the study was supposedly warning about.
It is a wonderful statistic. Shocking, institutionally badged, and short enough to survive a thousand reposts. It also does not say what almost everyone thinks it says.
The MIT NANDA report behind it did not find that enterprise AI was failing across the board. The same document reports that over 80% of organisations had explored or piloted general-purpose LLM tools, nearly 40% reported deployment, and generic LLM chatbots showed a pilot-to-implementation rate of around 83%. What actually struggled was custom, embedded, workflow-specific GenAI measured against a demanding six-month P&L test — a window that, in enterprise software, is short enough to make almost anything look like a failure. The authors themselves described their figures as directionally accurate and based on interviews rather than company reporting, and listed their limitations openly. Other large surveys from the same period found substantially more positive pictures.
So the headline number is not evidence that AI products fail. It is evidence that a specific kind of AI product, judged on a specific short clock, struggles to show returns.
That distinction is the entire subject of this article. Because the way most teams build AI roadmaps is: absorb a loud signal, misread what it measures, and ship against it. The 95% figure is just the most visible example of a much more common failure — mistaking noise for demand.
The signals most AI roadmaps are built on
Walk through the inputs behind a typical AI product roadmap and you'll usually find some combination of these:
A competitor shipped something. The most reliably destructive input in product management, and AI has made it worse, because everyone is shipping constantly and nobody knows which of it is working.
A model got better. A new release makes something possible that wasn't. The feature gets built because it's now buildable, not because anyone asked for it.
A customer said it would be cool. Not that they'd pay for it. Not that they'd change their workflow for it. They said it would be cool, in a call, while being polite.
A conference talk. Somebody's architecture diagram looked impressive on a slide.
A number like the one above. A statistic that confirms an existing instinct, absorbed without checking what it measured.
Every one of these is a real observation. None of them is a market signal, because none tells you that a specific person will change what they currently do in order to use what you build.
That gap shows up later as a cancelled project. Gartner predicts that over 40% of agentic AI projects will be cancelled by the end of 2027, and the reasons it gives are worth reading carefully: escalating costs, unclear business value, and inadequate risk controls. Not model capability. Their analyst's phrasing is blunter than most vendor commentary — most agentic AI projects are early-stage experiments driven by hype and often misapplied, and many use cases positioned as agentic today don't require agentic implementations.
The same research contains the detail that should worry anyone building in this space: of the thousands of vendors marketing agentic AI capabilities, Gartner estimated only around 130 were building the real thing. The rest were rebranding chatbots, assistants and RPA. The industry now calls it agent washing.
Which means a meaningful share of the competitive activity you're benchmarking against is not a signal about the market. It's a signal about somebody's marketing department.
What a real signal looks like
The useful test is simple: a real signal is evidence that someone is already paying a cost you could remove.
Not that they'd like a thing. That they are, right now, spending money, time, headcount or risk on a problem — and that the spending is visible to you.
Signals that pass the test:
Somebody is doing it manually, at volume, on purpose. The strongest signal in existence. A person is being paid to do a repetitive thing, which means the problem is real, the value is quantifiable, and the budget already exists. You aren't creating a line item; you're moving one.
Somebody built an ugly internal workaround. A spreadsheet with twelve tabs. A Zapier chain nobody dares touch. A contractor who does it every Friday. Workarounds are the clearest evidence of unmet demand there is, because someone has already paid to solve it badly.
Somebody is paying for an inadequate tool. Existing spend on a product they complain about is better evidence than enthusiasm for a product that doesn't exist. Budget is already allocated; only the recipient is in question.
Somebody churned, and told you why. Cancellation reasons are the least flattering and most honest data you will ever receive.
The same question arrives repeatedly, unprompted, in support. Not a feature request — a question. Questions reveal where your product's model of the world and the user's diverge.
Signals that fail the test:
- "That would be amazing" in a discovery call
- Waitlist signups with no payment step
- Competitor launches
- LinkedIn engagement
- Anything a model can now do that nobody asked for
- Analyst predictions about a market three years out
The distinction isn't optimism versus pessimism. It's evidence of existing cost versus expression of hypothetical interest. Only the first predicts behaviour.
The specific trap in AI products: capability is not demand
Traditional product development has a natural brake on it. Building something takes long enough that somebody usually asks why during the build.
AI removes the brake. You can ship an impressive demo of almost anything in a fortnight. Which means the question "does anyone need this?" gets skipped, because the cost of skipping it feels low — right up until the point where the demo has to survive real inputs, real volume, real edge cases, and a real P&L review.
Three specific ways this plays out:
Capability-led roadmaps. A model gains a capability, so the roadmap gains a feature. The logic runs backwards from what's possible instead of forwards from what's needed. The tell is a roadmap that reshuffles every time a model ships.
Novelty as differentiation. Being first to add an AI feature is a differentiator for roughly one quarter. Then everyone has it, and you're competing on the thing you were competing on before, having spent a quarter not improving it. Novelty is rented, not owned.
Demo-to-production collapse. The demo works on the ten inputs someone tested by hand. Production brings malformed data, ambiguous requests, timeouts and adversarial users. Anyone who has taken an AI feature to production knows the demo is roughly 10% of the work. Roadmaps built on demo velocity systematically underestimate the remaining 90%, which is where the escalating costs in Gartner's cancellation list actually come from.
There's a related trap worth naming. When a model's capability is your product's core value, your differentiation belongs to whoever trained the model — and it can be commoditised by the next release. Durable AI products are usually built where the model is one component and the moat is somewhere else: proprietary data, workflow integration depth, distribution, domain-specific evaluation, or accumulated feedback that makes the system better in ways a competitor can't copy by switching models.
Ask the uncomfortable version: if the model got twice as good and half as expensive tomorrow, would that help us or destroy us? If it destroys you, you're a feature of the model, not a product.
Finding the gap that's actually a gap
Competitive analysis in AI usually produces a feature matrix, which is close to useless. Everyone claims everything, and — per the agent-washing numbers — a lot of them are claiming things they don't have.
More useful questions:
What do their customers complain about, specifically? Review sites, support forums, churn commentary. Complaints are more reliable than feature lists because nobody writes a fake complaint about accuracy on a workflow they don't use.
What do they refuse to do? Constraints reveal architecture. A vendor who won't self-host, won't touch regulated data, won't integrate with a legacy system, or won't work below a certain contract size has drawn a boundary. Boundaries are where underserved customers live.
Who are they not built for? Most AI products are built for the well-resourced case: clean data, technical buyer, existing team to operate it. The unserved market is usually the same problem in a messier context — no data team, legacy stack, smaller budget, tighter regulation.
What are they measuring? If every competitor optimises for the same visible metric, the gap is often in the metric nobody reports. In support automation, everyone publishes deflection; almost nobody publishes true resolution or repeat-contact rate. In outbound, everyone publishes meetings booked; few publish meeting-to-opportunity conversion. The unmeasured metric is frequently where the real customer pain sits, precisely because no competitor is being held accountable for it.
What's genuinely hard? If a thing is easy, being good at it is not a position. If it's hard — messy integration, regulated environments, verification, unglamorous edge cases — competence is durable, because the difficulty deters entrants.
That last one deserves emphasis. Gaps that persist are usually gaps that are boring or difficult. The exciting gaps get filled within months by twelve funded startups.
Building the roadmap around it
Concretely, what changes.
- Every item carries its evidence. Beside each roadmap entry: what observed behaviour justifies it, and how confident you are. Items whose evidence reads "competitor has it" or "the model can now do this" get flagged. Not necessarily cut — flagged, so the ratio is visible. If most of the roadmap is capability-led, that's a strategy problem, and it's better seen than felt.
- Sequence by evidence strength, not excitement. Strongest evidence first, even when it's less interesting. The unglamorous item with three customers already paying to solve it manually beats the exciting item with enthusiastic verbal interest. Reliably.
- Define the falsifiable bet. Each significant item gets a written statement of what you believe, what would prove you wrong, and by when. Written before the build, not after. This is the single cheapest discipline available, and the one most consistently skipped.
- Budget the last 90%. Plan the demo and the production work: evaluation, error handling, monitoring, edge cases, cost control. If that isn't budgeted, you're planning a demo and calling it a product. Escalating cost is the leading named cause of agentic project cancellation, and it escalates because it was never estimated.
- Set a kill criterion before you start. Not a review date — a specific condition under which you stop. Teams are far better at killing projects when the criterion was agreed while everyone was still calm.
- Reserve capacity for what you'll learn. A roadmap with no slack cannot respond to a real signal when one arrives. If everything is committed, the only signals you can act on are the ones that arrived before the planning meeting.
- Judge on the right clock. The lesson of the 95% number is not that AI fails. It's that a six-month P&L window measures something narrower than success. Pick your measurement window deliberately, state it upfront, and don't let a leading indicator get evaluated as though it were a lagging one.
The uncomfortable part
Applied honestly, this process kills things you were excited about. That's the point. The roadmap that survives it is shorter, less impressive in a board deck, and considerably more likely to result in a product people pay for.
It also means saying no to the feature your competitor just launched, on the grounds that you have no evidence anyone wanted it and neither, probably, did they.
The AI market is going through a correction that was entirely predictable, and the projects that die in it will mostly not die of bad technology. They'll die of unclear business value — which is a polite way of saying nobody checked whether anyone needed the thing before building it.
The check is not expensive. It is about four questions: who is already paying a cost here? What are they paying it with? What would they have to stop doing to use what we build? And how would we know if we were wrong?
Most roadmaps can't answer them. The ones that can tend to be right.
Where to start this week
Take your current roadmap. Beside each item, write the specific observed behaviour that justifies it — an actual thing an actual person does today, not a stated preference.
Anything you can't fill in is a hypothesis. Keep them if you want; hypotheses are legitimate. But mark them, count them, and if they outnumber the evidence-backed items, you now know something useful about how your roadmap was built.
Sources: MIT NANDA, "The State of AI in Business 2025"; Gartner press release, 25 June 2025, on agentic AI project cancellations.
Working through this on your own roadmap? Get in touch.



