#AI#AI Adoption

Your AI Pilots Are Not Failing for the Reason You Think

MIT study that found 95% of AI initiatives fail

Prabhu Jay

Every AI conversation in a boardroom right now starts the same way. Someone mentions the MIT study that found 95% of AI initiatives fail, the room goes quiet, and the CFO asks why the company is on track to spend more on AI next year than it did on its last ERP upgrade.

That is a fair question. It is also the wrong first question. Very few of the people quoting the 95% have read where it came from, and almost none have looked at what it excludes. The number is real enough to take seriously and soft enough that using it to justify a spending freeze would be a mistake.

Here is what the research actually supports, and what we think leadership teams should do about it.

Where the number comes from

The source is a July 2025 report from MIT's Project NANDA called The GenAI Divide: State of AI in Business 2025. It looked at roughly 300 public AI deployments, ran 52 structured executive interviews, and surveyed about 150 leaders and 350 employees. Against $30 to $40 billion in enterprise GenAI investment, it found that 95% of pilots produced no measurable P&L impact and that only about 5% of custom-built enterprise AI tools ever reached production.

Two findings inside the report got far less attention than the headline, and both are more useful. Buying from specialized vendors succeeded roughly 67% of the time, while internal builds succeeded about a third as often. And mid-market companies moved from pilot to production in about 90 days, where large enterprises took nine months or longer. If you are a Fortune 500 CIO, that second finding should sting more than the first.

Now the caveats, because they matter. The report calls itself preliminary. It defined success as production deployment with measurable KPIs and P&L impact inside roughly six months. Wharton's Kevin Werbach read it several times, could not trace how the 95% was derived, and publicly asked MIT to release the data or pull the report. NANDA originated at MIT but is not administered by the university, and the report goes on to recommend NANDA's own agentic protocols as the fix. None of that makes the finding worthless. It does mean that anyone treating it as a precise measurement is being sold something.

Apply that six month test to the last two decades of enterprise technology and cloud migration fails, data warehousing fails, and most ERP programs fail spectacularly.

Why it still holds up

What makes the direction credible is not MIT. It is that four other research programs, with different methods and much larger samples, land in the same place.

McKinsey surveyed nearly 2,000 executives across 105 countries in 2025 and found 88% of organizations using AI in at least one function, but only 39% reporting any EBIT impact at all, and about 6% attributing more than 5% of EBIT to AI. BCG surveyed 1,250 senior leaders and put 5% in the top "future-built" category, 35% scaling, and 60% seeing almost nothing. Deloitte's 2026 survey of 3,235 leaders across 24 countries found two thirds reporting productivity gains but only a fifth reporting revenue growth, and only a quarter saying initiatives hit the ROI executives expected. S&P Global found 42% of companies abandoned most of their AI initiatives in 2025, up from 17% the year before. Gartner expects more than 40% of agentic AI projects to be canceled by the end of 2027.

Adoption is near universal. Value capture is in the single digits. The exact percentage is an argument for analysts. The gap is not.

The cost model most business cases get wrong

Sit through enough AI funding reviews and you notice the same omission. The business case covers model spend, cloud, and platform tooling, and stops there. Licensing and model fees are frequently a minority of what the thing actually costs to run.

What shows up later is data remediation, which is almost always the largest single effort line. Then integration into the systems that matter, meaning ERP, CRM, claims, or the EMR. Then evaluation, testing, and red teaming. Then governance and monitoring, which now consumes 8 to 12% of the average AI budget against 3 to 5% in 2024. Then training and process redesign. Then the maintenance nobody budgets for, because model versions get deprecated and prompts rot.

The FinOps Foundation found 73% of enterprises exceeded their original AI cost projections in 2026. Other 2026 survey work put the average overrun near 31%, with a quarter of organizations off by half or more.

The reason forecasting keeps failing is worth understanding, because it is counterintuitive. Unit prices are collapsing. One analysis of 2.4 billion enterprise API calls showed blended cost per million tokens falling 67% year over year. Invoices went up anyway, because consumption grew faster than price fell. Agentic workflows burn an order of magnitude more tokens per task than a chatbot does, and the spread between the cheapest production model and a frontier reasoning model runs past 4,000x. Organizations that route workloads by task landed near $2.31 per million tokens. Organizations that send everything to the best available model paid closer to $18.40 for the same work.

Put simply, this is a usage-based cost sitting inside a company that budgets like it buys seats. Uber gave 5,000 engineers access to an AI coding tool in December 2025 and burned the entire annual AI budget by April. That is not a failure of the technology.

Why programs stall

Across all of the research, the failures are not model quality failures. They cluster into a handful of patterns that any experienced delivery leader will recognize from earlier technology cycles.

The biggest is that nothing about the work changes. AI gets bolted onto a process that is left intact, which produces a good demo and no P&L movement. McKinsey found the highest EBIT impact among organizations redesigning end-to-end workflows. BCG puts a number on it with what they call 10-20-70: about 10% of the value comes from the algorithm, 20% from data and technology, and 70% from people and process. Most budgets are allocated in almost exactly the reverse proportion.

Second, pilots run on clean curated data and production runs on the real thing, which is siloed, inconsistent, and poorly governed. That gap is where timelines double.

Third, the build reflex. Most enterprises still default to building, particularly in regulated industries, despite the evidence that vendor partnerships convert at roughly twice the rate.

Fourth, value gets defined after the fact. Gartner attributes agentic cancellations to escalating costs, unclear business value, and weak risk controls, in that order. Only 36% of finance chiefs in Gartner's 2026 CFO survey said they were confident in their ability to deliver real enterprise impact from AI.

Fifth, and this one is almost cultural, boards fund what they can see. Customer-facing AI demos well. The MIT case studies found the strongest returns sitting in the back office, where nobody presents them: $2 to $10 million in annual savings from displacing outsourced support and document review, and roughly 30% cuts in external agency spend.

There is one more signal that should reframe the whole conversation. In more than 90% of firms, employees are already using personal AI tools while the sanctioned pilot goes nowhere. The demand is not the problem.

Where the value actually is

The upside is real, but it is concentrated in a way that should change how you allocate.

IDC's research for Microsoft puts the average return at $3.70 per dollar invested, with leading adopters near $10 and median time to positive ROI around 13 to 14 months. BCG's future-built 5% report 1.7x the revenue growth, 3.6x the three-year shareholder return, 2.7x the return on invested capital, and 1.6x the EBIT margin of laggards. BCG also found roughly 70% of AI value sitting in core functions such as R&D, sales and marketing, and operations rather than in the adjacent experiments that are easier to get approved.

That points to a simple allocation rule. A small number of deeply integrated use cases inside work that already matters will beat broad tool distribution every time, and it is not close.

When we build a value case with a client, four things have to be measurable before funding: cost per transaction against a real baseline, revenue effect attributable to the specific step AI touches, error and rework rates including the new AI-specific controls, and capacity with an explicit plan for where the freed capacity goes. The most common way a good case turns into a bad outcome is counting hours saved that never leave the cost base.

If a case cannot survive a 14 month payback and a 30% cost overrun, it is not ready to fund.

How to run the portfolio

The practical fix is not more governance. It is separating three kinds of work that most companies currently run under one set of expectations.

TierWhat it coversHow to fund itHow you judge it
EfficiencyBack office, document, support workflowsFull funding against a hard cost baselineCost per transaction down, vendor spend actually removed
GrowthCustomer and revenue-facing workflowsStaged funding against leading indicatorsConversion or cycle time improvement that holds two quarters
LearningCapability building, emerging techCapped and time-boxed, no ROI testA documented decision to kill or scale inside 90 days

Most of the reported failure rate is Tier 3 work being measured with Tier 1 expectations. Fix that and the numbers look different immediately.

Beyond the tiering, four things separate programs that ship from programs that get canceled. Every initiative has a named business owner who controls the P&L line it claims to move. Buy is the default for commodity capability, partner for domain depth, build only where the process is genuinely proprietary. Baselines get captured before the pilot starts, not reconstructed afterward to justify it. And cost attribution runs at the workflow level from day one, with engineering and finance meeting monthly, because the architectural choices engineers make determine most of the invoice that finance sees six weeks later.

If you want a starting sequence: spend the first month inventorying every active AI initiative and killing anything without an owner or a baseline. Spend the second picking two workflows with hard cost baselines and running data readiness on both. Spend the third shipping one of them into real production with a real owner, and publish the result whether or not it worked.

What to ask in the next review

The 95% figure is not evidence that AI does not work. It is evidence that most organizations are running AI as a technology program when the things going wrong are operational, financial, and organizational.

Three questions will tell you where you stand faster than any maturity assessment. For each funded initiative, which P&L line does it move and who owns that line. What does the cost look like at three times current usage. And what work are we redesigning rather than decorating.

Companies that cannot answer those are not investing in AI. They are funding pilots and hoping.


Sources: MIT Project NANDA, The GenAI Divide: State of AI in Business 2025 (July 2025) and published methodology critiques; McKinsey State of AI 2025 (n=1,993); BCG The Widening AI Value Gap: Build for the Future 2025 (n=1,250) and AI Radar 2026; Deloitte State of AI in the Enterprise 2026 (n=3,235); Gartner worldwide AI spending forecast 2026, agentic AI cancellation forecast, and 2026 CFO priorities survey; IDC and Microsoft The Business Opportunity of AI; S&P Global; FinOps Foundation State of FinOps 2026.

0claps
P

Prabhu Jay

Head of Engineering