AI exposes your enterprise before it improves it

Why enterprise AI pilots quietly stall, and a simple test to find where yours will break next.

Share
The word Legibility set large on The Beagle’s green field — AI at Scale.

You have sat in this meeting. The quarterly AI review. The slide says the pilot was a success. And it was, a year ago, in a demo, when everything went right. The number on it hasn’t moved since. Around the table, everyone keeps circling the same question, each in their own words. Why isn’t this becoming an advantage?

The answer never makes it onto a slide. The AI worked exactly as designed. It didn’t fail your business. It reflected it.

We’ve been sold one story about AI. Point it at your company and it will transform it. Reengineer the process. Unlock the value. The promise is always the same: transformation, something done to the enterprise, from the outside, by the technology.

That story is backwards. AI doesn’t arrive as an engine that reshapes whatever it touches. It arrives as a mirror.

The machine can only run what you wrote down

The mechanism is simple, which is why the same thing happens in every industry. An AI system can only act on what your business has made explicit: the rules it can read, the definitions it can find, the decisions someone wrote down. Everything else it guesses, or ignores. And most of what runs your business was never written down.

Take one word. Customer. Sales means the entity that signs. Finance means the billing account. Support means the person who calls. Three definitions, three systems, all correct, none lined up. For years, people have quietly bridged that gap without noticing they were doing it. An AI can’t. It picks one, silently, and scales the mismatch to ten thousand cases a day.

One word, three definitions
“Customer”
Sales
the entity that signs
Finance
the billing account
Support
the person who calls
Three definitions, all correct, none lined up. People bridge the gap without noticing. An AI can’t — it picks one, silently, and scales the mismatch to ten thousand cases a day.

Why your pilots don’t add up

Each one rebuilds its own private, partial version of the business: its own guess at what a customer is, how a price is set, when an exception applies. None of it adds up to a capability. It piles up as interpretations, each one slightly different, owned by no one.

Your org chart says someone owns “data.” Nobody owns meaning. (There is rarely a committee for that.)

By some estimates, more than 80% of AI projects fail, about twice the rate of IT projects that don’t involve AI. Treat that as a direction, not a decimal: it’s a figure people quote, not one anyone has cleanly measured. What sits underneath it is firmer. When RAND asked sixty-five senior data scientists why projects die, the reasons sorted into five. Four are about the organisation: the business never agreed on the problem, the data wasn’t there, the team chased the tool instead of the outcome, or nothing was built to run in real use. Only the fifth is a true technical limit — the task was simply too hard for today’s models. Four failures out of five happen before the model is even the question.

Why AI projects die — RAND’s five root causes
Business never agreed on the problem
The data wasn’t there
Chased the tool, not the outcome
Nothing built to run in real use
Task too hard for today’s models
Organisational · 4 of 5 Technical · 1 of 5
Four of the five root causes are organisational; only one is technical. The often-quoted “over 80% of AI projects fail, about twice the IT rate” is an estimate RAND cites, not a number it measured. Source: RAND, The Root Causes of Failure for AI Projects, 2024.

There’s a word for this, and it isn’t a technical one

The word is legibility, and legibility, not your ambition, is what AI scales. It means how easily a system can be read from the outside, by someone who did not grow up inside it. The political scientist James C. Scott named it, describing states trying to govern societies they could not see. The parallel is close enough to be useful. Your AI is a central authority trying to act on a business it cannot yet read.

So the question was never which model, which platform, which vendor. It is quieter than that, and harder. How legible is your business to itself?

Point a capable model at a business whose logic is clear, consistent, and governed, and it builds on itself. Point it at one held together by expert memory and local workarounds, and it spreads that confusion faster than any person can catch.

So a stalled AI programme isn’t the technology’s fault. It’s a reading of how clearly your business understands itself. The model is the thermometer, not the fever.

None of this means transformation is impossible, or that you must map the entire enterprise before you touch AI. That’s its own kind of failure, and an expensive one. The smaller, more useful claim is this. AI will expose your business before it improves it, and the exposure is the point. It shows you exactly where to make the business explicit. Not everywhere. Only where it keeps breaking.

Run the mirror on yourself this week

You can do this with no budget and no vendor. Pick one decision that matters. How a discount gets approved. When a claim is escalated. What makes an account “at risk.” Ask five people who own that decision to write down, in one sentence, how it actually gets made. Then count the different answers.

One answer, and the decision is legible. AI can scale it safely. Five answers, and you’ve just found where your next pilot will quietly break, and what to fix before you build on top of it.

The legibility test
  1. Pick one decision that matters — how a discount is approved, when a claim is escalated, what makes an account “at risk”.
  2. Ask five people who own it to write, in one sentence, how it actually gets made.
  3. Count the distinct answers.
One answer = legible; AI can scale it safely. Five = where your next pilot breaks. Do it for ten decisions and you are holding a map of where your business can’t yet see itself.

Do it for ten decisions and you’re holding a map. Not of your technology. Of the places your business can’t yet see itself.

What fixing one square is worth

Say the discount decision came back with five answers. The fix isn’t a data-lake programme. It’s smaller and less glamorous than that: the five owners sit in one room and agree, in writing, on a single definition and the three conditions that change it. A week of work, not a quarter.

Written down and shared, that agreed definition is the discipline behind legibility — the first line of a business ontology, the explicit layer that decides whether AI compounds your advantage or scales your contradictions. And what happens next is far bigger than the effort. Every future pilot that touches pricing now inherits one rule instead of guessing at five. The next AI project starts from a definition it can read, not a contradiction it has to solve. Each decision you make legible lowers the cost of the one after it. That is how the advantage builds.

This is also where control quietly sits. A definition you own and govern stays on your side of the table. One you leave unspoken gets decided by whatever rented model reaches it first, on rules you did not set. The moat was never the model. Anyone can rent that.

The moat is the stock of decisions your competitors still can’t state in a single sentence.

The companies that pull ahead in the next phase won’t be the ones with the best models. They’ll be the ones who looked first, and had the discipline to fix what it showed them.

The mirror is already up in your organisation. The only question left is whether you’re reading it, or still arguing with the slide.

Common questions

Why do AI pilots fail to scale?

Usually not because the model is weak. When RAND studied why AI projects fail, four of the five root causes were organisational — the business never agreed on the problem, the data wasn’t there, the team chased the tool, or nothing was built to run in real use. A pilot scales whatever the business already is; where the business contradicts itself, the pilot stalls.

What is the legibility test?

Pick one decision that matters, ask five people who own it to write in one sentence how it actually gets made, and count the different answers. One answer means the decision is legible and AI can scale it safely. Five answers mark exactly where your next pilot will break.

Is stalled AI a technology problem?

Rarely. A stalled AI programme is usually a reading of how clearly your business understands itself — the model is the thermometer, not the fever. The fix is to make the business explicit where it keeps breaking, not to buy a bigger model.