In August 2025, a research group at MIT published a number that has become the most quoted figure in enterprise software. Roughly ninety-five percent of generative AI pilots were failing to deliver measurable P&L impact. The work was based on 150 leader interviews, a 350-person employee survey, and a review of 300 public deployments. The number landed, traveled, and is now used by every consulting firm in the country, including the firms that sold the failed pilots.
Most leaders read the number and conclude that the technology isn’t ready, or their organization isn’t ready, or the pilots needed more time. That framing is incomplete, and the data doesn’t support it. The mechanism is different.
The 5 percent that worked share a structural posture toward scope and accountability. A model, a vendor, or an architecture in common would have been easier to copy. The posture is teachable, most enterprise AI programs skip it, and it explains the gap better than any technical variable.
The number is doing different work than people think
Read carefully, the MIT NANDA finding is a verdict on procurement. The technology mostly did what it was asked. The successful 5 percent were disproportionately external partnerships rather than internal builds. They started with a single workflow, not a platform, and wrote down what success would look like in dollars before the contract was signed.
Compare that to the prevailing pattern. McKinsey’s State of AI Trust in 2026 reports that organizations putting twenty-five million dollars or more into responsible AI initiatives report higher maturity scores. The framing is that responsible AI is a spending category. It can be. But it can also be a workflow contract with one number on it.
Gartner’s CIO Agenda 2026 runs the same arithmetic from another angle. Ninety-four percent of CIOs expect major plan changes inside twenty-four months. Only forty-eight percent of digital initiatives meet or exceed business targets. The fail rate is the long-running base rate for IT projects that begin without a measurable outcome attached. It is neither novel nor unique to AI.
That’s the structural read. The 95 percent is a story about how mid-market and enterprise leaders contract for software that’s supposed to do work.
The takeaway: the 95 percent is the old IT project failure rate under a new name, and the shared defect is a contract with no measurable outcome attached. Before blaming the model, go find the number the pilot promised to move.
What the successful 5 percent actually do
Three things, repeatedly. None of them are exotic.
First, they pick a workflow that has a number on it. A function is not a workflow, and neither is a department or a “use case.” Accounts receivable past forty-five days qualifies. So do renewal-quote turnaround and inbound support resolution time. The success criterion is the change in that number, measured against a baseline taken before the work begins.
Second, they cap the pilot at a length where outcomes can be observed. Six to twelve weeks is long enough to install and short enough that no one can hide. Gartner’s research on agentic AI governance notes that organizations with formal governance platforms are roughly 3.4 times more likely to achieve high effectiveness. The platforms are an artifact of the leaders who already wrote down what they were measuring, not the cause.
Third, they negotiate the engagement with the outcome embedded in the price. If the workflow does not move, the bill changes. This is the part that exposes the rest of the program. Vendors who wouldn’t sign that contract told the buyer, in advance, that they didn’t believe their own pitch.
Together those three moves separate the 5 percent from the 95 percent. They are procurement and governance decisions, not technology decisions. The technology is downstream of the contract.
The takeaway: the 5 percent won before the contract was signed, with one numbered workflow and a price tied to the outcome. Put those terms in front of your vendor and see who still wants the deal.
Why most leaders get this wrong
Three patterns recur.
The first is that AI is treated as a horizontal capability rather than a workflow intervention. The board hears “we are deploying AI across the company” and approves a budget. There is no single workflow on the hook for results. Twelve months later, no single workflow has changed enough to defend the spend. McKinsey’s State of Organizations 2026 reports that eighty-eight percent of organizations experiment with AI but eighty-one percent report no meaningful bottom-line impact. That’s a scoping problem. The experimentation was never the issue.
The second is that internal builds win the political argument inside IT and lose the economic one outside it. The MIT data is consistent on this. Buying from a specialist vendor and forming a partnership succeeds at roughly twice the rate of internal builds. The internal build feels safer because it preserves control. It is more expensive and slower, and less accountable. The mid-market in particular cannot afford that posture. A $90M industrial distributor does not have the engineering bench to compete with a focused vendor on a niche workflow. It can buy clarity faster than it can build it.
The third is that pilots are scoped to prove the technology rather than to change the workflow. A pilot designed to demonstrate that an LLM can read invoices is not the same as a pilot designed to reduce DSO by ten days. The first ends with a slide deck, the second with cash. Most programs run the first kind, then ask why the board is unimpressed.
CIO Magazine in 2026 reported on a related dynamic: agentic AI systems do not fail catastrophically. They drift. Behavior changes incrementally as models update, prompts evolve, and tool integrations are added. A workflow that worked in week six can be worse in week thirty. The drift only matters if someone is still measuring. Most pilots stop measuring after the press release.
The takeaway: most failed programs scoped AI company-wide and then graded the pilots on activity instead of outcomes. Put one workflow on the hook for a dollar number and keep measuring after the headline.
The actual mechanism
The mechanism that separates signal from noise in enterprise AI is a small number of structural decisions made before the contract. The model barely enters into it.
Pick one workflow, define the baseline, cap the engagement, tie price to the outcome, and keep measuring after the headline.
Each of those is boring. None of them require a Fortune 50-caliber data science team, and all of them are within reach of a $25M to $250M company. That is the upside in the failure data. To win on AI, the mid-market needs to be more disciplined than the average enterprise on scope and accountability, which is a much easier bar to clear than competing with hyperscalers on infrastructure.
The mid-market is where this mechanism runs cleanest. A $90M industrial distributor with one ERP, one CRM, and roughly fifteen senior leaders can pick a single workflow, name a baseline, and produce a measurement cadence inside a quarter. A Fortune 100 enterprise running twelve business units cannot. The friction at scale is political. Every workflow has three claimants and four dotted-line owners. Naming a single baseline becomes a negotiation. The 95 percent failure rate is partly an artifact of running the playbook at the size where the playbook is hardest to run. That is also why the mid-market wins faster on AI when it bothers to be disciplined. The size is the advantage.
What changes the picture
A mid-market CIO who is running a half-dozen AI pilots in May 2026 has a small number of moves available before the next budget cycle.
Pilot Audit
- Which workflow does this change?
- What is the dollar baseline?
- What does the contract say happens if the dollar number does not move?
Pilots that cannot answer all three get paused. Pilots that can get funded harder.
Treat AI governance as a measurement cadence, not a platform purchase. The billion-dollar AI governance market Gartner is forecasting will sell tools to companies that needed a measurement habit. The habit is the governance. The tool can be helpful only after the habit exists.
Resist the “agentic” framing as a budget category. Agentic systems are workflow tools. They drift, and they require a person who owns the outcome and watches the number. The fix lives in the procurement contract.
The 95 percent is a story about contracts that shouldn’t have been signed and pilots graded on activity instead of outcomes. Read that way, the number is useful. It points directly at the move, and the move is procurement discipline.
That’s the entire game in mid-market AI right now.
The Yale Chief Executive Leadership Institute, working with Sonnenfeld and colleagues, used the release of Anthropic’s most powerful model in early May 2026 as a forcing function for a governance framework across banking, healthcare, retail, and supply chain. The framework is useful. It is also telling that it took a model release to get boards to write down what they were measuring. The companies that already had a governance posture did not need the model release. They had the posture because they had the contracts, the contracts had the numbers, and the numbers were the governance.
A note on implementation: procurement teams are not, by training or incentive, the right owners of an outcomes contract on AI. Procurement optimizes for unit price. Outcomes contracts optimize for value capture, which often costs more on a sticker basis. The CIO who delegates this to a procurement-led RFP is buying the wrong thing for the wrong reason. The contracts that produce the 5 percent are negotiated by an operating leader who can name the workflow, the baseline, and the dollar number, then handed to procurement for terms. Reverse that sequence and the result is the 95 percent.