Skip to content

Why AI Pilots Fail: The Failure Is Built Into the Format

Why AI Pilots Fail: The Failure Is Built Into the Format

A pilot has a property nobody says out loud: by definition, it isn't obligated to change anything. It has no process owner, no integration with the system of record, and no date after which the old way of working shuts down. Its only job is to prove the technology works.

And the technology almost always works. That's exactly why the conversation keeps drifting back to models and vendors, even though the gap between a successful outcome and a failed one isn't created there. Most pilots fail by design, not by content.

In short

A pilot tests the technology. The decision to deploy it requires answers to three different questions: who owns the process after launch, where does the output land, and when does the old route get shut down. The pilot answers none of them — and isn't supposed to. The fix is simple: stop building pilots to demonstrate capability and start building them to measure one number — the share of outputs that need correction.

35%of executives don't know who has the authority to override an AI recommendation
34%have no mechanism for resolving conflict between a person and the model
15%of organizations are mature in both AI adoption and change management

The three questions a pilot never answers

Let's separate what a pilot actually tests from what it leaves untested. The list is short, and every item on it decides the fate of the rollout more than the quality of the model's answers does.

Who owns the process after launch. During the pilot, the project team owns it, and that team disbands once the pilot ends. What follows is that the process has no one accountable for the error rate, no one authorized to halt the scenario, and no one who signs off on rule changes. Without that person, the scenario survives only until its first serious failure.

Where the output lands. During the pilot, only the participants see the result — that's enough to judge quality. In production, the output has to land in the system where the company's facts of record live, or no one outside the original request will ever see it. Integration is usually left out of pilot scope because it costs more than the pilot itself.

When the old route gets shut down. Running a pilot alongside existing work is normal — that's how pilots are supposed to operate. But parallel operation doesn't end on its own: as long as the old way stays available at no cost, part of the workload keeps flowing through it, and separating out the effect becomes impossible.

QuestionWhat the pilot testsWhat stays untested
Output qualityFully tested
Process ownerNot testedWho is accountable for the error rate and for stopping the scenario
Landing in the system of recordNot testedWhether anyone outside the original request ever sees the output
Sunset of the old routeNot testedWhether the workload migrates in full
Operating costPartially testedSupport, updates, error triage

Notice that the pilot does exactly the job it was assigned, flawlessly. The problem isn't the pilot — it's that people read its result as an answer to the deployment question.

What the data show about undefined AI decision rights

A missing owner isn't an abstraction — it's a measured fact, measured across a large sample of executives.

Research

35% of surveyed executives said it's still unclear who has the authority to challenge or override an AI recommendation, and 34% have no durable mechanism for resolving conflict between a human decision and a model's output. Only 15% of organizations have reached maturity in both AI adoption and change management at the same time.

IBM Institute for Business Value with Oxford Economics, June 2026. Survey of 1,000 senior executives responsible for AI, data, and technology across 14 countries and 21 industries, quota sample screened for genuine involvement; a separate survey of 8,400 employees across 28 countries · ibm.com

This is a sponsored study from a consulting arm whose goal is to sell an operating-model redesign, and the data are correlational, not causal. That's a real limitation. But the number itself describes a fact, not a judgment call: a third of executives can't name who has the authority to override the system's decision. That is exactly the question a pilot never asks.

A pilot answers "does the technology work." Failure comes from the question "who is accountable for it."

Why AI pilot gains disappear after launch

There's a body of work that explains the mechanism behind gains earned through a one-off push evaporating later — and it predates the current wave of technology by decades.

Research

Improvements achieved through a burst of intensified attention, without being embedded into an ongoing process, systematically reverse: once the metrics recover, people get pulled off that work, and the problems come back. The authors call this the capability trap.

Nelson P. Repenning, John D. Sterman (MIT Sloan). "Nobody Ever Gets Credit for Fixing Problems that Never Happened," California Management Review, Summer 2001. More than a decade of field research, over 12 in-depth case studies across industries, system dynamics modeling, interviews, and archival analysis · web.mit.edu (PDF)

This work has nothing to do with AI and was written a quarter-century ago — it's an analogy, not direct proof, and I'm using it as one. Still, the mechanism it describes is precise, and it's recognizable in every pilot: three months of intensified team attention produce a result that gets attributed to the technology rather than the attention, and it disappears along with it.

There's a second side to this mechanism too — a systematic overstatement of results from the mere fact of being observed.

Research

A re-analysis of recovered original data from the famous illumination experiments found that the textbook picture of a sharp productivity jump from the mere fact of being observed doesn't hold up: "the existing descriptions of ostensibly notable data patterns turn out to be entirely fictional," though weak signs of an observation effect are present.

Steven D. Levitt, John A. List (University of Chicago). NBER Working Paper 15016, May 2009; later published in American Economic Journal: Applied Economics, 2011. Search for and recovery of original data from the 1920s experiments, thought lost, found on microfilm, followed by formal quantitative reanalysis · nber.org

The full text is paywalled, and I worked from the abstract and the summary of findings — that's a limitation. The practical takeaway cuts both ways, which is what makes it useful: you can't credit "the Hawthorne effect" for every pilot success — it's weaker than conventional wisdom suggests. But you also can't carry pilot results into production without adjusting for the team's heightened attention — the adjustment is just smaller than folklore claims.

When to kill an AI pilot versus when to fix the threshold

It's worth dropping a false frame here too: the fact that most experiments never reach production isn't, by itself, a failure.

Research

Machine learning engineers describe the path from experiment to production as a selection process: it pays off for organizations to prototype many ideas quickly and cull them before the best one reaches production use. The widely cited share of models that never make it to production reflects the nature of experimentation, not a failure.

Shreya Shankar, Rolando Garcia, Joseph M. Hellerstein, Aditya G. Parameswaran (UC Berkeley). "Operationalizing Machine Learning: An Interview Study," arXiv, September 2022. Semi-structured interviews with 18 machine learning engineers from companies of varying size and industry, coded using grounded theory · arxiv.org

The sample is small — 18 people — and not statistically representative; the authors say so themselves. But their observation defuses a false alarm: the problem isn't that nine pilots out of ten get shut down. It's that the cutoff criterion was never set in advance. When there's no criterion, the scenarios that get killed aren't the worst ones — they're the ones whose sponsor lost interest.

How to design an AI pilot that predicts production results

You don't need to cancel the pilot — you need to change the question it answers. Instead of "does the technology work," it should answer "what share of outputs will need correction."

A pilot with a threshold set in advance: one measured number, one cutoff, one decision at the end.
01What we measure: the share of outputs that need correction
02The threshold at which the scenario moves forward
03Measurement on the real workflow, not on cherry-picked examples
04The decision: deploy, redesign, or kill
the threshold and the kill criterion are set before the pilot starts, not after the results come in

The second node is the key one. A threshold named in advance turns the pilot from a demonstration into a test, and turns any dispute over the result into a check against a number. Without it, the outcome gets argued over by impression, and the decision goes to whoever argues most persuasively.

The third node matters just as much. A real workflow differs from cherry-picked examples in exactly the cases that justify deploying the technology in the first place: nonstandard phrasing, incomplete data, edge cases. A demonstration on prepared data tells you nothing about them. For how to turn these same numbers into a business case, see this separate breakdown.

Four decisions to make before you start an AI pilot

01

Name the process owner, not the project owner

The person who stays on after the pilot ends and is accountable for the error rate. Without that person, the pilot is pointless.

02

Set the acceptance threshold before you start

What share of outputs requiring correction counts as acceptable. Name the number in advance and write it down.

03

Test on the real workflow

No pre-screening of examples. It's the nonstandard cases that determine whether the scenario is fit for use.

04

Name the date the old route gets sunset

Not inside the pilot itself, but in the deployment plan — and it has to be named before the pilot starts, or the deployment never finishes.

A pilot without a threshold set in advance isn't a test of a hypothesis. It's a search for arguments to support a decision that's already been made.

Frequently asked questions about why AI pilots fail

Why don't AI pilots make it to production deployment?

Because a pilot tests the quality of the technology, while moving to production requires answers to three other questions: who owns the process, where does the output land, and when does the old way of working get shut down. None of them are part of the pilot format. The pilot ends successfully and then runs into decisions nobody prepared for.

How should you set up a pilot for an AI use case?

Reduce it to measuring one number — the share of outputs that will need correction — and set an acceptance threshold before you start. Take the measurement on the real workflow, without pre-selecting examples. What comes out is a number, not an impression, and the decision gets made by comparing it to the threshold, not by arguing about it.

Should you kill a pilot if the result is ambiguous?

Yes, and a kill criterion set in advance is the only way to do it without conflict. Culling most experiments is normal — it's a property of experimentation, not a sign of failure. What's abnormal is when the scenarios that get shut down aren't the worst ones but the ones whose sponsor lost interest — that's how an organization loses its ability to learn from its own attempts.

How well do pilot results carry over into production?

With an adjustment for the team's heightened attention, which disappears in production. That adjustment is smaller than the popular version of the observation effect suggests — a re-analysis of the classic experiments found it to be modest. But it isn't zero: a pilot that runs on the sponsor's daily hands-on involvement will noticeably underperform in production.

Where the internal figures come from

The list of three questions a pilot never answers, and the pilot format built around a threshold set in advance, are an original synthesis drawn from technology-project evaluation practice across the China-Russia Investment Fund's portfolio and from deployment practice at Alego.Digital. No internal numeric metrics are cited in this article: every figure comes from an external source with its method stated. No formal write-up of this format exists in company documents; it's presented here for the first time.

External sources
  1. IBM Institute for Business Value, Oxford Economics. The AI-Human Operating Model, June 2026 (survey of 1,000 executives across 14 countries). ibm.com
  2. Repenning N. P., Sterman J. D. Nobody Ever Gets Credit for Fixing Problems that Never Happened. California Management Review, 2001. web.mit.edu
  3. Levitt S. D., List J. A. Was There Really a Hawthorne Effect at the Hawthorne Plant? NBER Working Paper 15016, 2009. nber.org
  4. Shankar S., Garcia R., Hellerstein J. M., Parameswaran A. G. Operationalizing Machine Learning: An Interview Study. arXiv:2209.09125, 2022. arxiv.org