AI exposes your enterprise before it improves it

The pilot did exactly what it was built to do. What it reflected back is the problem, and a test that costs you about an hour shows you where.

Share
The word Legibility set large on The Beagle’s green field — AI at Scale.

You have sat in this meeting. The quarterly AI review, the slide that says the pilot was a success, and the number on that slide, which has not moved in a year. Around the table the question goes round in different words, and it is always the same question: why is this not becoming an advantage? By the end of this you will be able to answer it for your own company, with one test that takes about an hour and no vendor, and to know which decisions to make explicit first, which is where an AI advantage starts to build.

The honest answer never makes it onto the slide, because it flatters nobody in the room. The AI did exactly what it was built to do. It did not fail your business. It reflected it.

You have been sold a story about AI, and so have I. Point it at your company and it will transform it: reengineer the process, unlock the value, something done to the enterprise from the outside, by the technology. That story is backwards. AI does not arrive as an engine that reshapes whatever it touches. It arrives as a mirror. So what exactly does it see when it looks at your business?

What the machine sees, and what it cannot

An AI system can only act on what your business has made explicit: the rules it can read, the definitions it can find, the decisions somebody took the trouble to write down. Everything else it guesses or ignores, and most of what runs a business was never written down at all. Ask yourself how much of your own department lives in three people's heads.

Take one word, customer. For Sales it is the entity that signs, for Finance it is the billing account, for Support it is the person who calls. Three definitions, three systems, all of them correct and none of them lined up. For years, people have bridged that gap without noticing, the way you step over a loose tile in your own kitchen. An AI cannot. It picks one definition, silently, and scales the mismatch to ten thousand cases a day.

One word, three definitions
"Customer"
Sales
the entity that signs
Finance
the billing account
Support
the person who calls
Three definitions, all correct, none lined up. People bridge the gap without noticing. An AI can't. It picks one, silently, and scales the mismatch to ten thousand cases a day.

That gives you a better sentence for the next review than "the model is wrong". Try "the model picked one of our three customers, and nobody told it which". Now multiply that by every pilot you have running. What do they add up to?

Why your pilots do not add up

Each pilot rebuilds its own private, partial version of the business: its own guess at what a customer is, how a price is set, when an exception applies. None of it adds up to a capability; it piles up as interpretations, each slightly different from the next and owned by nobody. Some of that starts at selection, since nobody chose these pilots from a list of AI use cases expecting them to add up. Your organisation chart says someone owns "data", and nobody owns meaning.

When RAND asked sixty-five senior data scientists and engineers why projects die, the reasons sorted into five. Four are about the organisation: the business never agreed on the problem, the data was not there, the team chased the tool instead of the outcome, or nothing was built to run in real use. Only the fifth is a true technical limit, the task being too hard for today's models. Four of the five causes sit before the model is even the question.

RAND stops there, and so should I: it does not claim those four organisational causes share a single root. That connection is my argument, not RAND's finding. Still, if four of the five causes sit with the organisation, it is fair to ask what they have in common, and my answer comes from an unlikely place.

Why AI projects die: RAND's five root causes
Business never agreed on the problem
The data wasn't there
Chased the tool, not the outcome
Nothing built to run in real use
Task too hard for today's models
Organisational · 4 of 5 Technical · 1 of 5
RAND lists five root causes; four sit with the organisation. The split into organisational and technical is my reading, not a label RAND uses. Source: RAND, The Root Causes of Failure for AI Projects, 2024.

The word for it, and where it comes from

The word is legibility. Legible means readable from the outside, by someone who did not grow up inside; a business is legible when a newcomer, or a machine, can find its rules without asking anyone. The word comes from James C. Scott's book Seeing Like a State (1998): a state cannot tax or count people it cannot see, so states spent centuries making people visible, with fixed family names, one unit of measure and maps of who owns what.

Your AI is that newcomer. It is trying to act on a business it cannot yet read, and legibility, not your ambition, is what it scales. Point a capable model at a business whose logic is clear, consistent and governed, and it builds on itself. Point it at one held together by expert memory and local workarounds, and it spreads that confusion faster than any person can catch it. In short, a stalled AI programme is a reading of how clearly your business understands itself, not a verdict on the technology. The model is the thermometer, not the fever.

So the question was never which model, which platform, which vendor. It is a quieter question and a harder one: how legible is your business to itself? You cannot answer it by mapping the whole enterprise first; that is its own kind of failure, and an expensive one. The states took centuries. You have until the next review, which is where the hour comes in.

The one-hour test

For my part, I have stopped asking a management team what their rule is, because every answer is sincere and they never match. Instead, pick one decision that matters: how a discount gets approved, when a claim is escalated, what makes an account "at risk". Ask five people who own that decision to write down, in one sentence, how it is really made. Then count the different answers. This is what I call the Legibility Test, and it costs no budget, no vendor and about an hour of your time, plus a day for the answers to come back.

The legibility test
  1. Pick one decision that matters. How a discount is approved, when a claim is escalated, what makes an account "at risk."
  2. Ask five people who own it to write, in one sentence, how it really gets made.
  3. Count the distinct answers.
One answer = legible; AI can scale it safely. Five = where your next pilot breaks. Do it for ten decisions and you are holding a map of where your business can't yet see itself.

One answer, and the decision is legible, so AI can scale it safely. Five answers, and you have just found where your next pilot will break, and what to fix before you build on top of it. Do it for ten decisions and you are holding a map, not of your technology but of the places where your business cannot yet see itself.

What fixing one square is worth

Say the discount decision came back with five answers. The fix is smaller and less glamorous than a data-lake programme: the five owners sit in one room and agree, in writing, on a single definition and the three conditions that change it. A week of work, not a quarter. The meeting is not comfortable, because somebody in that room discovers that their rule was never the rule. But it is short, and it is a safe bet that it will be the most useful week your AI programme has had.

Written down and shared, that agreed definition is the discipline behind legibility. It is the first line of a business ontology, the explicit layer that decides whether AI builds on your advantage or multiplies your contradictions. Every future pilot that touches pricing now inherits one rule instead of guessing at five. Each decision you make legible lowers the cost of the one after it. That is how the advantage builds: from the stock of decisions you can state in one sentence and your competitors cannot, not from the model.

This is also where control sits: a definition you own and govern stays on your side of the table, and one you leave unspoken gets decided by whatever rented model reaches it first. The moat was never the model.

The moat is the stock of decisions your competitors still cannot state in a single sentence.

A written definition decays as soon as the business moves on, unless somebody owns it. So each one gets an owner, and a reason to be reopened: a new product, a new market, a pilot that breaks on it. The review is the same room, an hour, not a committee. What I have not seen settled is how light that can stay in a large group before it turns into paperwork nobody follows; if you have kept it light, I would like to know how.

The rule, then, is simple: the companies that pull ahead in the next phase will not be the ones with the best models. They will be the ones who looked first, and had the discipline to fix what the mirror showed them. So this week, before the next review, take one decision that matters and run the test: five owners, one sentence each, about an hour, then count. And you: how many answers do you think you will get?

Common questions

Why do AI pilots fail to scale?

Usually not because the model is weak. When RAND studied why AI projects fail, four of the five root causes sit with the organisation: the business never agreed on the problem, the data wasn't there, the team chased the tool, or nothing was built to run in real use. A pilot scales whatever the business already is; where the business contradicts itself, the pilot stalls.

What is the legibility test?

Pick one decision that matters, ask five people who own it to write in one sentence how the decision gets made in practice, and count the different answers. One answer means the decision is legible and AI can scale it safely. Five answers mark exactly where your next pilot will break.

Is stalled AI a technology problem?

Rarely. A stalled AI programme is usually a reading of how clearly your business understands itself. The model is the thermometer, not the fever. The fix is to make the business explicit where it keeps breaking, not to buy a bigger model.