Somewhere in your company right now, there’s a Slack channel for an AI pilot that used to be exciting. It has a name like #ai-pilot or #genai-taskforce. Six months ago it was full of screenshots and “this could change everything” energy. Today it’s quiet. The pilot technically still exists. Nobody’s turned it off. It just… didn’t go anywhere.
If that sounds familiar, you’re not doing anything wrong. You’re doing what almost everyone is doing with AI pilots. MIT’s NANDA initiative found that roughly 95% of enterprise generative AI pilots fail to produce measurable ROI.
Not “underperform.” Fail to show up in the numbers at all.
We heard an almost identical story from operators at an industry conference for Registered Investment Advisors (RIAs), independently of the MIT research: AI pilots stall out for the same three reasons, again and again: data too fragmented to trust, workflows that span too many people and systems for a point tool to touch, and compliance requirements nobody accounted for until the pilot was already live.
Here’s the part that matters most for you as an operator: The models are good enough; what is failing is everything around the model. The bad news is that’s usually where nobody was looking, because everyone was busy evaluating vendors.
We’ve sat inside enough of these AI pilots to see the pattern clearly enough to name it. And naming it is useful, because once you can see the pattern, you can see your own way out of it.
The symptom you can see: point-solution sprawl
Before we get to the framework, it’s worth describing what “AI adoption outrunning AI readiness” actually looks like from a COO’s chair, because it rarely looks like failure. It looks like activity.
Copilot is live in one department. Agentforce is being piloted in another. Half your managers have ChatGPT open in a browser tab they use for things they’d never put in an official system. Everyone is using AI. Nobody would say your company has “adopted AI” in any way that shows up on a P&L, because none of these tools talk to each other, none of them touch the actual handoffs between people and systems, and none of them were designed around how work really moves through your org.
This is what happens when AI gets bolted onto a process instead of the process getting rebuilt around AI. And that distinction (bolting on vs. redesigning around) turns out to be the single biggest predictor of whether an AI pilot becomes durable production or becomes a quiet Slack channel.
A framework for where you actually are
When we work through this with operators, we use a four-stage maturity model, because “are we good at AI or not” is too blunt a question to act on. The useful question is which stage of the work is actually broken.
-
Foundational. AI use is ad hoc and individual. People use consumer tools on their own initiative. There’s no shared data foundation, no governance, and no organizational memory of what’s working. Every “win” lives in one person’s workflow and dies when they change roles.
-
Developing. Pilots exist. Departments have picked tools. There’s some executive attention. But integration is shallow. Each tool is still an island, and most of what’s automated is a single step inside a much longer process, not the process itself.
-
Advancing. AI touches real, cross-functional workflows. Data foundations are solid enough to trust. There’s a governance layer that isn’t just a policy document nobody reads. Outcomes are being measured, not just usage.
-
Transforming. AI is a structural part of how the operation runs, not an add-on to it. Workflows were rebuilt with AI as a given, not retrofitted around a legacy process. This is rare, and it should be. It’s not where every function needs to sit, and treating it as a scoreboard to “win” misses the point.
Most companies we talk to are somewhere between Foundational and Developing with their AI pilots, and most of them believe they’re further along than that, because activity feels like progress. It isn’t the same thing.
Wave 1 vs. Wave 2: why AI pilots stall or scale
The maturity model tells you where you are. This next distinction tells you why you might be stuck there.
Wave 1 is AI layered onto an existing process without changing the process. You keep the same approval chain, the same handoffs, the same systems that don’t talk to each other, and you insert an AI step somewhere inside it. It’s the fastest way to get a pilot running, which is exactly why almost everyone starts here. It’s also why almost everyone stalls here: you’ve made one step faster while the rest of the workflow, and all its fragmentation, stays exactly as broken as it was before.
Wave 2 is redesigning the workflow around what AI actually makes possible. Instead of asking “where can we drop a chatbot into this process,” you ask “if we were building this process today, knowing what AI can do, what would it look like?” That’s a harder, slower, more uncomfortable question, and it’s the one the 5% are asking.
This is the real difference between the AI pilots that plateau and the ones that become production. It’s not model quality. It’s not budget. It’s whether the workflow itself was ever actually rethought.
What the 5% of AI pilots that reach production do differently
A few patterns show up consistently in the organizations that get past the AI pilot stage:
They fix the data problem before the AI problem. Fragmented, siloed data was the number-one blocker operators named independently of each other. If your AI can’t see across the systems the workflow actually touches, it can only ever automate a fragment of the work, which is what Wave 1 looks like from the inside. (We wrote about what it takes to give AI that context in How Does a Semantic Layer Unlock Your Data for Agentic AI.)
They design for the whole workflow, not the task. A task is “summarize this document.” A workflow is “how a claim moves from intake to approval across four people and three systems.” The 5% start by mapping the second thing, not the first.
They build governance in from the start, not bolted on after. Compliance was the third blocker operators named, and it’s almost always because governance gets treated as a later step instead of a design constraint from day one. Retrofitting compliance onto a live pilot is far more painful than designing for it upfront.
They measure outcomes, not activity. “We deployed a tool” isn’t a metric. “We cut internal search time by 80%” is, which is roughly what Swirl AI Connect delivered once AI was wired into how their teams actually search internal knowledge, instead of sitting next to it as a separate tool. Or take a company like Knapsack, running AI agents in actual production, inside SOC 2, HIPAA, and GDPR constraints, which only works because governance was part of the workflow design, not an afterthought bolted on once something broke.
None of this requires a moonshot. It requires being honest about which stage you’re actually in, and being willing to touch the workflow itself instead of decorating it with a new tool. For a broader look at the path from assessment to production, read From legacy systems to production AI.
Where to start
If you’re reading this and recognizing your own Slack channel, the useful next move is an honest diagnosis: which of the seven dimensions (data, workflow design, governance, and the rest) is actually the thing holding you at Foundational or Developing, and are you closer to Wave 1 than you’d like to admit?
We built a free assessment that walks through exactly this. 24 questions, about 10–12 minutes, scored across the same seven dimensions and the same maturity model above, so you get a real read on where you sit rather than a vibe.
You can run it here: cheesecakelabs.com/ai-readiness-assessment.
No pitch required to use it. But if what comes back points at a workflow that needs to be rebuilt rather than another tool bolted on, that’s the conversation worth having next.