The Human and Machine Company Free → Read it
← Writing

Your AI Pilot Succeeded. So What.

The pilot worked. The demo was impressive. Leadership is excited. And in six months, nothing will have changed. The gap between a successful AI pilot and captured economic value is where most companies lose, and almost nobody is talking about why.

The demo went well. The pilot showed a 40 percent reduction in processing time, the team is energized, and leadership is cautiously optimistic. The steering committee wants to see a plan for scaling.

Most companies call this a success. It’s the moment I start to worry.

I’ve watched enough of these play out to know what comes next. The scaling plan gets drafted but not funded. The team that ran the pilot moves to other priorities. The vendor contract renews on autopilot. Six months later, the tool is technically still live but nobody is measuring whether it does anything. The pilot succeeded. The company didn’t capture a dollar.


The pilot-to-production gap is not a technology problem

RAND Corporation research published in 2025 dissected the failure modes of enterprise AI. The results are specific: 80 percent of projects fail to deliver promised business value. Of those, a third are abandoned before production. Another 28 percent reach production but fail to deliver expected value. Eighteen percent run indefinitely without recovering their investment.

80%of AI projects fail to deliver promised business valueRAND Corporation, 2025

That last category is the one that should concern you most. These are the zombies. They are live, they are consuming budget, and nobody is asking whether the number moved. In the mid-market, where budgets are tighter and every dollar has a name, zombie initiatives are particularly expensive because they consume the organizational attention that should be directed at the next opportunity.

IDC’s research quantified the conversion rate: for every 33 AI proofs of concept an enterprise starts, only four reach production. That is a 12 percent graduation rate. Most of the proofs of concept work technically. The constraint is everything that happens after they work.

The takeaway: only four of every 33 proofs of concept reach production, and the rest either die or run on as unmeasured zombies. Go find the zombies in your own budget before you start pilot number 34.


What “success” actually requires

A pilot proves that a technology can do something. It doesn’t prove the something is worth doing at scale, that the organization can absorb the change, or that anyone will bother to measure the outcome.

The companies that actually capture value from AI pilots do three things that most companies skip.

They name the dollar baseline before the pilot starts. A projected savings estimate doesn’t count, and neither does a vague efficiency metric. The baseline is a specific number in a specific workflow that exists today, before the AI touches anything. Days sales outstanding. Renewal-quote turnaround time. Support resolution time. Cost per processed claim. If you cannot name the number before the pilot, you will not be able to prove the pilot changed it.

A pilot without a kill date is a subscription.

They set a kill date. A pilot without a kill date is a subscription. The kill date forces a measurement conversation. Either the number moved or it did not. The pilot graduates to production with a real budget, or it stops. S&P Global found that the share of enterprises abandoning most AI initiatives jumped from 17 percent in 2024 to 42 percent in 2025. The companies that abandoned were finally measuring, not failing faster.

They assign a P&L owner, not a project manager. Project managers run pilots, and value capture needs an owner whose compensation is tied to the number. This distinction matters because scaling an AI tool is mostly change management, workflow redesign, and financial measurement. A project manager keeps the timeline, and a P&L owner keeps the accountability.

The takeaway: a real pilot has a dollar baseline named up front, a kill date on the calendar, and an owner paid on the number. Refuse to launch anything missing one of the three.


The gap between “it works” and “it’s worth it”

In my diagnostic framework, I score companies across five dimensions. The lowest-scoring dimension, consistently, is whether AI results show up in the financials. That means the ability to convert AI activity into dollars the business can see, measure, and act on.

A successful pilot that never shows up in the P&L is like a profitable product line with no accounting entry. The value exists in theory, but nobody can find it in the numbers. It doesn’t inform capital allocation, show up in board materials, or change how the company thinks about its next investment.

The Consero Global 2026 CFO Survey found that fully embedded AI in finance functions grew 91 percent year over year. But Bain & Company’s survey of senior finance executives tells the rest of the story: only 15 to 25 percent of CFOs have fully scaled AI in their departments, and satisfaction is sharply higher among those who have (41 percent) compared to those still running pilots (25 percent). The gap is completion, not enthusiasm.


Pilot purgatory is a structural choice

The mid-market version of pilot purgatory has a specific shape. A company runs a pilot. The pilot works. The team requests budget to scale. The budget request goes to a planning cycle that meets quarterly. By the time the budget is approved, the original team has moved to other priorities, the vendor has released a new version, and the business case needs to be re-argued from scratch.

I have seen this cycle repeat three and four times at the same company, with different tools, different teams, and the same outcome, so the tools are off the hook. The company has no mechanism for converting a pilot result into a funded production program within the window where the result is still valid.

The fix is structural. Before the pilot starts, decide what happens if it works. Specifically: who approves the production budget, what the approval threshold is, and how long the decision takes. If you can’t answer those three questions before the pilot launches, you’re running a science fair project.

The takeaway: pilot purgatory is a choice, made every time a pilot launches with no plan for what happens if it works. Decide who approves production money, at what threshold, and on what clock, while the pilot is still a whiteboard sketch.


The question nobody asks

Here is the question I ask every prospective client, and it catches people off guard every time: “Of the AI initiatives you have completed in the past 18 months, how many have a dollar value attached to them that your finance team would stand behind?”

The answer is almost always zero or one. Occasionally two.

The initiatives didn’t fail. Many of them worked beautifully. Nobody built the measurement infrastructure alongside the technology. The pilot team measured accuracy, speed, and adoption. Nobody measured payback. The finance team was not in the room when the pilot was designed, so the financial framework was never built, and now it is too late to construct one retroactively because the baseline data was not captured.

Writer’s 2026 enterprise survey of 2,400 executives found that 79 percent report real friction in their AI programs and only 29 percent report meaningful returns. The gap between “we are using AI” and “AI is making us money” opens at the pilot stage, not at the scaling stage. By the time you are trying to scale, the measurement opportunity has already passed.

The takeaway: the measurement window opens at the pilot stage and closes before scaling, because a baseline nobody captured cannot be rebuilt later. Put your finance team in the room on day one of pilot design.


What to do before your next pilot launches

Audit every current AI initiative against one question: does this initiative have a named dollar number, a named owner, and a kill date? For every initiative that cannot answer all three, decide whether it is a measurement program or an activity program. Activity programs generate reports, and measurement programs generate P&L impact.

For your next pilot, write down the production decision framework before the pilot starts. Name who approves production funding, what threshold triggers approval, and how long the decision takes. If the framework does not exist, the pilot has no path to production regardless of how well the technology performs.

The Pilot Audit

A named dollar number, a named owner, a kill date. Who approves production funding and at what threshold. If those answers do not exist before the pilot starts, the pilot has no path to production regardless of how well the technology performs.


How RLK Can Help

My AI Diagnostic measures whether your AI results are showing up in the financials, alongside four other dimensions, and identifies exactly where the pilot-to-production gap sits in your organization. If the gap is structural, the AI Business Case engagement builds the financial framework, owner accountability, and decision timeline that turns a working pilot into a line item. Reach out.


Sources

Ryan King

About the author

Ryan King

Fifteen years in technology strategy at McKinsey and Deloitte. Now running RLK Consulting: enterprise-caliber tech strategy, one strategist doing every hour of the work. Over $10B in documented value capture across 50+ engagements and 12 industries.