Why AI Pilots Fail: The Failure Is Built Into the Format
A pilot has a property nobody says out loud: by definition, it isn't obligated to change anything. It has no process owner, no integration with the system of record, and no date after which the old way of working shuts down. Its only job is to prove the technology works.
And the technology almost always works. That's exactly why the conversation keeps drifting back to models and vendors, even though the gap between a successful outcome and a failed one isn't created there. Most pilots fail by design, not by content.
A pilot tests the technology. The decision to deploy it requires answers to three different questions: who owns the process after launch, where does the output land, and when does the old route get shut down. The pilot answers none of them — and isn't supposed to. The fix is simple: stop building pilots to demonstrate capability and start building them to measure one number — the share of outputs that need correction.
The three questions a pilot never answers
Let's separate what a pilot actually tests from what it leaves untested. The list is short, and every item on it decides the fate of the rollout more than the quality of the model's answers does.
Who owns the process after launch. During the pilot, the project team owns it, and that team disbands once the pilot ends. What follows is that the process has no one accountable for the error rate, no one authorized to halt the scenario, and no one who signs off on rule changes. Without that person, the scenario survives only until its first serious failure.
Where the output lands. During the pilot, only the participants see the result — that's enough to judge quality. In production, the output has to land in the system where the company's facts of record live, or no one outside the original request will ever see it. Integration is usually left out of pilot scope because it costs more than the pilot itself.
When the old route gets shut down. Running a pilot alongside existing work is normal — that's how pilots are supposed to operate. But parallel operation doesn't end on its own: as long as the old way stays available at no cost, part of the workload keeps flowing through it, and separating out the effect becomes impossible.
| Question | What the pilot tests | What stays untested |
|---|---|---|
| Output quality | Fully tested | — |
| Process owner | Not tested | Who is accountable for the error rate and for stopping the scenario |
| Landing in the system of record | Not tested | Whether anyone outside the original request ever sees the output |
| Sunset of the old route | Not tested | Whether the workload migrates in full |
| Operating cost | Partially tested | Support, updates, error triage |
Notice that the pilot does exactly the job it was assigned, flawlessly. The problem isn't the pilot — it's that people read its result as an answer to the deployment question.
What the data show about undefined AI decision rights
A missing owner isn't an abstraction — it's a measured fact, measured across a large sample of executives.
35% of surveyed executives said it's still unclear who has the authority to challenge or override an AI recommendation, and 34% have no durable mechanism for resolving conflict between a human decision and a model's output. Only 15% of organizations have reached maturity in both AI adoption and change management at the same time.
This is a sponsored study from a consulting arm whose goal is to sell an operating-model redesign, and the data are correlational, not causal. That's a real limitation. But the number itself describes a fact, not a judgment call: a third of executives can't name who has the authority to override the system's decision. That is exactly the question a pilot never asks.
Why AI pilot gains disappear after launch
There's a body of work that explains the mechanism behind gains earned through a one-off push evaporating later — and it predates the current wave of technology by decades.
Improvements achieved through a burst of intensified attention, without being embedded into an ongoing process, systematically reverse: once the metrics recover, people get pulled off that work, and the problems come back. The authors call this the capability trap.
This work has nothing to do with AI and was written a quarter-century ago — it's an analogy, not direct proof, and I'm using it as one. Still, the mechanism it describes is precise, and it's recognizable in every pilot: three months of intensified team attention produce a result that gets attributed to the technology rather than the attention, and it disappears along with it.
There's a second side to this mechanism too — a systematic overstatement of results from the mere fact of being observed.
A re-analysis of recovered original data from the famous illumination experiments found that the textbook picture of a sharp productivity jump from the mere fact of being observed doesn't hold up: "the existing descriptions of ostensibly notable data patterns turn out to be entirely fictional," though weak signs of an observation effect are present.
The full text is paywalled, and I worked from the abstract and the summary of findings — that's a limitation. The practical takeaway cuts both ways, which is what makes it useful: you can't credit "the Hawthorne effect" for every pilot success — it's weaker than conventional wisdom suggests. But you also can't carry pilot results into production without adjusting for the team's heightened attention — the adjustment is just smaller than folklore claims.
When to kill an AI pilot versus when to fix the threshold
It's worth dropping a false frame here too: the fact that most experiments never reach production isn't, by itself, a failure.
Machine learning engineers describe the path from experiment to production as a selection process: it pays off for organizations to prototype many ideas quickly and cull them before the best one reaches production use. The widely cited share of models that never make it to production reflects the nature of experimentation, not a failure.
The sample is small — 18 people — and not statistically representative; the authors say so themselves. But their observation defuses a false alarm: the problem isn't that nine pilots out of ten get shut down. It's that the cutoff criterion was never set in advance. When there's no criterion, the scenarios that get killed aren't the worst ones — they're the ones whose sponsor lost interest.
How to design an AI pilot that predicts production results
You don't need to cancel the pilot — you need to change the question it answers. Instead of "does the technology work," it should answer "what share of outputs will need correction."
The second node is the key one. A threshold named in advance turns the pilot from a demonstration into a test, and turns any dispute over the result into a check against a number. Without it, the outcome gets argued over by impression, and the decision goes to whoever argues most persuasively.
The third node matters just as much. A real workflow differs from cherry-picked examples in exactly the cases that justify deploying the technology in the first place: nonstandard phrasing, incomplete data, edge cases. A demonstration on prepared data tells you nothing about them. For how to turn these same numbers into a business case, see this separate breakdown.
Four decisions to make before you start an AI pilot
Name the process owner, not the project owner
The person who stays on after the pilot ends and is accountable for the error rate. Without that person, the pilot is pointless.
Set the acceptance threshold before you start
What share of outputs requiring correction counts as acceptable. Name the number in advance and write it down.
Test on the real workflow
No pre-screening of examples. It's the nonstandard cases that determine whether the scenario is fit for use.
Name the date the old route gets sunset
Not inside the pilot itself, but in the deployment plan — and it has to be named before the pilot starts, or the deployment never finishes.
A pilot without a threshold set in advance isn't a test of a hypothesis. It's a search for arguments to support a decision that's already been made.
Frequently asked questions about why AI pilots fail
Why don't AI pilots make it to production deployment?
Because a pilot tests the quality of the technology, while moving to production requires answers to three other questions: who owns the process, where does the output land, and when does the old way of working get shut down. None of them are part of the pilot format. The pilot ends successfully and then runs into decisions nobody prepared for.
How should you set up a pilot for an AI use case?
Reduce it to measuring one number — the share of outputs that will need correction — and set an acceptance threshold before you start. Take the measurement on the real workflow, without pre-selecting examples. What comes out is a number, not an impression, and the decision gets made by comparing it to the threshold, not by arguing about it.
Should you kill a pilot if the result is ambiguous?
Yes, and a kill criterion set in advance is the only way to do it without conflict. Culling most experiments is normal — it's a property of experimentation, not a sign of failure. What's abnormal is when the scenarios that get shut down aren't the worst ones but the ones whose sponsor lost interest — that's how an organization loses its ability to learn from its own attempts.
How well do pilot results carry over into production?
With an adjustment for the team's heightened attention, which disappears in production. That adjustment is smaller than the popular version of the observation effect suggests — a re-analysis of the classic experiments found it to be modest. But it isn't zero: a pilot that runs on the sponsor's daily hands-on involvement will noticeably underperform in production.
The list of three questions a pilot never answers, and the pilot format built around a threshold set in advance, are an original synthesis drawn from technology-project evaluation practice across the China-Russia Investment Fund's portfolio and from deployment practice at Alego.Digital. No internal numeric metrics are cited in this article: every figure comes from an external source with its method stated. No formal write-up of this format exists in company documents; it's presented here for the first time.
- IBM Institute for Business Value, Oxford Economics. The AI-Human Operating Model, June 2026 (survey of 1,000 executives across 14 countries). ibm.com
- Repenning N. P., Sterman J. D. Nobody Ever Gets Credit for Fixing Problems that Never Happened. California Management Review, 2001. web.mit.edu
- Levitt S. D., List J. A. Was There Really a Hawthorne Effect at the Hawthorne Plant? NBER Working Paper 15016, 2009. nber.org
- Shankar S., Garcia R., Hellerstein J. M., Parameswaran A. G. Operationalizing Machine Learning: An Interview Study. arXiv:2209.09125, 2022. arxiv.org