Alert: Beware of Scams! We've been informed of fraudulent schemes using Netsurit's name. As a rapidly growing brand, we recognize the threat of cybercriminals exploiting our reputation.Cybercrime is on the rise, and we are committed to protecting our community.
Why AI Pilots Fail

Why AI Pilots Fail: What the 95% Statistic Really Means

By Dr Jan van Niekerk, SVP Data & AI Strategy, Netsurit  |  6 August 2026

Most AI pilots labelled “failures” did not fail to work. They failed a strict test: no measurable P&L impact within six months, which is exactly what the widely quoted 95% statistic measures. In August 2025, a Fortune headline did the rounds of every board pack I saw that quarter: “95% of generative AI pilots at companies are failing.” I watched a steering committee quote it as the reason to freeze a programme that was, by its own numbers, working.

The stat was real. The conclusion drawn from it was wrong. That gap, between a true number and a false conclusion, is what this first issue is about.

What the Study Actually Measured

The number comes from MIT’s Project NANDA, and the methodology deserves more attention than the headline got. The team reviewed more than 300 AI initiatives, interviewed people at 52 organisations, and surveyed 153 senior leaders. The famous 95% refers to custom, workflow-specific generative AI tools that showed no measurable P&L impact within six months.

Read that definition slowly. Custom builds. Six months. Visible in the P&L.

That is a strict test, and most transformation spending of any kind would fail it. Hold your ERP programme, your culture initiative, or your last reorganisation to “move the P&L within six months, measurably” and watch what happens to their success rates. Several analysts made this point within weeks of publication. The more interesting story was sitting in the authors’ own data.

Where the Value Actually Went

The same MIT report found that roughly 90% of employees use personal AI tools daily for work, while only about 40% of companies hold official enterprise subscriptions. The researchers called it a shadow AI economy.

Value was flowing the whole time. It was simply flowing through personal accounts, unmeasured and unsanctioned, while the official pilot with the steering committee and the six-month clock stalled in review cycles.

There is a second buried finding. Externally bought tools succeeded about 67% of the time in MIT’s data. Internal builds succeeded half as often. Put those together and the failure headline becomes a story about how enterprises buy, build, and measure, and much less a story about what the technology can do.

Three Studies, One Number

If MIT’s method bothers you, use someone else’s. BCG surveyed more than 1,250 decision-makers in 2025 and found 4 to 5% of companies capturing substantial value, with that group pulling away from the pack at roughly 1.7 times the revenue growth of peers. McKinsey’s 2025 State of AI (n=1,993) found 88% adoption, 39% reporting any EBIT impact, and 5.5% attributing more than 5% of EBIT to AI.

Three teams. Three methods. One answer, within a rounding error. About one organisation in twenty is getting real money out of AI, while nearly all of them are using it.

When independent methodologies converge like that, the message is structural. Adoption stopped being the constraint. Converting adoption into money is the constraint now.

What the Converters Do Differently

Across the AI programmes I’ve been close to, on four continents and in a dozen industries, the ones that crossed into production shared four habits. None of them involve model selection.

  • They baseline before they pilot. If nobody measured the process before the pilot, the after is theatre.
  • They measure with a stopwatch, not a survey. METR’s 2025 randomised experiment is the cautionary tale here. Experienced developers using AI tools were 19% slower on real tasks, while estimating afterwards that they had been 20% faster. Perception is not a KPI.
  • One name owns the outcome. Not a committee, not a working group. A name.
  • The kill date exists before the kick-off. Deciding the ending upfront turns a failed pilot from an embarrassment into a planned outcome, which, paradoxically, is what lets teams move quickly.

Honesty requires one concession to the pessimists. Some pilots deserve to die, and the 40-plus percent of agentic AI projects that Gartner expects to be cancelled by 2027 will not all be measurement victims. I can’t always tell in advance which is which, and I don’t trust anyone who claims they can. That is exactly why the baseline and the kill date matter. They let reality make the call, quickly and cheaply.

So before you approve the next pilot, ask for three things in writing. The number it is supposed to move, the date by which it must move, and the name of the person who signs when it doesn’t.

The next time someone quotes the 95% at you in a meeting, ask which 95% they mean. The pilots that failed a six-month P&L test, or the organisations that never defined the test at all.

Frequently Asked Questions

1. Why do most AI pilots fail?

Most don’t fail to function. They fail a strict measurement test: no measurable P&L impact within six months, using custom, workflow-specific tools. That’s the definition behind the widely quoted 95% figure. Many of those tools were still being used; they just weren’t tracked against that bar.

2. What is the actual AI pilot success rate?

Independent surveys converge on a similar number from different angles: MIT, BCG, and McKinsey each found that roughly 5% of organisations are converting AI adoption into measurable profit impact, even though adoption itself is now close to universal.

3. How long should an AI pilot run before you judge it?

Six months is a reasonable outer limit, not because of one paper but because any longer without a decision point invites indefinite pilot purgatory. What matters more than the ceiling is having a “before” baseline: without one, no after-number is verifiable.

4. Who should own an AI pilot’s outcome?

One named individual, not a committee. Pilots with a single accountable owner and a pre-agreed kill date are far more likely to either scale successfully or be shut down early, rather than drifting in review indefinitely.

5. Where This Leaves You

The 95% statistic isn’t wrong. It is answering a narrower question than most boardrooms think. The organisations pulling ahead are not the ones with better models. They are the ones who baseline first, measure outcomes instead of sentiment, name an owner, and set a kill date before day one.

If your board is still arguing about pilot counts, that is the wrong scoreboard. Read Netsurit’s whitepaper, Stop Counting Pilots, Start Counting Dollars, for how to turn AI productivity into a number your CFO will actually accept.

author avatar
Netsurit

Like this article?

Share on Facebook
Share on Twitter
Share on Linkedin
Share on WhatsApp
Share on E-mail