Skip to content
Back to blog

Industry Analysis

The Production Gap: Why So Few AI Pilots Reach Production

Alexander Snyder8 min

The number gets passed around boardrooms like a warning label: 95% of enterprise generative-AI pilots deliver no measurable P&L return. It comes from one source: MIT's Project NANDA, in a 2025 report called The GenAI Divide: State of AI in Business, based on 52 executive interviews, 153 survey responses, and roughly 300 publicly disclosed deployments.

It's worth knowing what that is. The report is preliminary by its authors' own description and has not been peer-reviewed. The sample is small, the measurement window is short, and the derivation of the 95% is not reconstructable from the published document, which is why academics have publicly asked MIT to release the underlying data. The figure also measures measurable P&L return inside the study window, which is not the same claim as "the pilot never shipped," even though that is how it usually gets repeated. Treat it as a directional signal from one non-peer-reviewed study rather than as a finding. It is directionally consistent with what I see in the field, and I would hold it loosely for exactly that reason.

Every AI vendor on the planet has a slide about it, usually right before they pitch their solution to the problem.

But here's what nobody talks about: the pilots aren't failing because the technology is bad. They're failing because the incentive structure guarantees failure.

The structural problem

Think about how most enterprise AI engagements work. A consultancy shows up with a team of four to six people. They spend two months in discovery. They build a proof of concept. They present results to the steering committee. They hand off a deck and a demo environment. They leave.

The client is now responsible for getting a prototype, built by people who don't understand the business, into a production environment that the prototype was never designed for. The consultancy has moved on to the next engagement. The internal team has their regular workload plus a new system to figure out.

The pilot dies. Not because it was technically flawed. Because nobody stuck around to finish the work.

Why the consultancy model breaks

The traditional consulting model optimizes for two things: billable hours and new logos. Neither incentivizes production deployment. A pilot that generates a great case study deck is just as valuable to the firm as a system that runs every day for years. Maybe more valuable, because the case study generates leads.

This creates a perverse incentive. The consultancy gets paid whether the system ships or not. The internal team gets blamed when it doesn't. And the next vendor pitch starts with "your last AI initiative failed because they didn't do it right."

The cycle repeats.

What production actually requires

Getting an AI system into production isn't a technology problem. It's an organizational problem. Here's what it actually takes:

Business context that takes months to build. You can't automate a process you don't understand. Understanding a process means attending the 7am calls, learning the names, sitting in on the meetings nobody wants to attend. There are no shortcuts.

Iteration tolerance. One engagement went through five pivots before we found the product that worked. Five. Most firms would have declared success after the first prototype and moved on. The client's patience, and our willingness to keep building, is the only reason that system is running today.

Failure documentation. When our dedup query was broken on a data enrichment engagement, we didn't bury it. We quantified the wasted spend down to the batch and presented it to the client. They expanded the engagement. Not despite the transparency. Because of it.

Ongoing presence. The best work happens after launch. Every system we've built has gotten better in the months after deployment. Optimization, expansion, new use cases. These only emerge when you're still there to see them.

The embed-or-fail hypothesis

We've started calling this the embed-or-fail hypothesis: if the team building the AI system isn't embedded in the organization, truly embedded, not just on-site for meetings, the system won't make it to production.

Every engagement we've run supports this. The ones that work are the ones where we joined the team. The ones that would have failed are the ones where we would have stayed outside.

What this means for buyers

If you're evaluating AI consultancies, ask one question: what happens after the pilot?

If the answer involves a handoff, a documentation package, and a knowledge transfer session, you're looking at a pilot factory. That's not inherently bad. But you should budget for the reality that production deployment will cost 3-5x the pilot, require different skills, and take significantly longer than anyone is telling you.

Or you can find someone who stays.


PurviewX is embedded AI leadership for companies sitting on real operational data. We find out whether your AI actually works, including ours. Start a conversation.