How to Run an AI Pilot That Reaches Production

95% of generative AI pilots deliver no measurable P&L impact, according to The GenAI Divide: State of AI in Business 2025, published by MIT's NANDA initiative and based on 150 interviews with business leaders, a survey of 350 employees, and analysis of 300 public AI deployments. For enterprise-grade custom tools the funnel is starker still: 60% of firms evaluated them, 20% reached pilot stage, and 5% went live.

Gartner reads the same landscape forward, predicting that over 40% of agentic AI projects will be cancelled by the end of 2027, and its named causes are worth reading twice: escalating costs, unclear business value, inadequate risk controls. Model capability is absent from the list. So is technical complexity.

Both studies reach the same diagnosis from different directions. The MIT researchers attribute the failure rate to a learning gap in how organisations integrate these tools, since executives tend to blame regulation or model performance while the evidence points at enterprise integration. Gartner's analyst calls most current projects early-stage experiments driven by hype and often misapplied. In plain terms: the pilots mostly work, and the projects around them mostly fail. Which means the difference between the 95% and the 5% is set at scoping, before any system is switched on.

The demo-shaped pilot

The common pilot is shaped like a demonstration. It selects the cleanest inputs available, because messy inputs would cloud the capability question. It recruits the most enthusiastic users, because sceptics slow things down. It measures accuracy against the old manual process, because that comparison flatters. And it reports to a steering group at the end, with a recommendation.

A point to note: Mitochondria goes for the hardest bits in a demo.

Each choice is reasonable for answering "can this technology do the task?" But that question is now largely settled, and the MIT data shows how thoroughly: employees in over 90% of surveyed firms already use personal AI tools for work, sanctioned or otherwise. The unsettled question is whether a system survives contact with the operation: the ordinary Tuesday, the malformed input, the user who never volunteered, the week the champion is on leave.

A demo-shaped pilot cannot answer that, because it was scoped to avoid encountering it. So it succeeds, the success contains no information about production, and the organisation discovers this at precisely the moment enthusiasm was supposed to convert into rollout. The pilot gets extended, revisited, or politely shelved, and joins the 95%.

MIT's most practical finding describes the mechanism: pilots stall because the tools cannot retain feedback, adapt to context, or improve over time. A static system demonstrated on static inputs looks finished. The same system exposed to a live operation, where the inputs drift and the exceptions compound, decays within weeks unless something in the design lets it learn the specific context it now works in. Demo-shaped pilots never test for this, because demos do not run long enough for decay to show.

The shift-shaped pilot

The pilot that converts is shaped like a shift. Five design choices, all made at scoping.

Run on ordinary inputs from day one. Whatever arrives, arrives: the scanned document at an angle, the supplier who fills the form in wrong, the request phrased three ways by three people. Performance on this material is the number that predicts production, and the only one worth collecting. A pilot that needs its inputs cleaned first has redefined the task into one the operation does not have.

Include the sceptic. Two or three users who did not ask for the system, alongside those who did. Their objections during the pilot are free consulting: each is either a design flaw caught early or an adoption obstacle mapped in advance. The same objections during rollout cost ten times more, because by then they have an audience.

Name an owner in the operating team before starting. A person whose function the work belongs to, who holds the system during the pilot and keeps it after. Gartner's finding that unclear business value drives cancellations is, at pilot scale, an ownership finding: value stays unclear when nobody whose budget it touches is responsible for demonstrating it. A pilot run by a vendor or an innovation team, to be handed over if successful, converts poorly, because the handover transfers a system without transferring the habit of running it.

Write exit criteria as a rate, before starting. Something of the form: over six weeks on live inputs, the system handles at least X% of cases without intervention, flags uncertain cases rather than guessing, and the owning team confirms the redesigned process runs. Numbers agreed after results are known are advocacy. Numbers agreed before are a decision rule, and a decision rule is what lets a pilot end in a decision rather than another meeting.

Rehearse the boring machinery. Access rights, escalation when the system is unsure, behaviour when it is down, who reviews the weekly numbers. Gartner's third cancellation cause, inadequate risk controls, is this machinery neglected. None of it is intellectually interesting, and all of it is what production consists of. A pilot that skips it has tested the model and postponed the system.

Six weeks, then a decision

Length matters less than a deadline. Six to eight weeks on live inputs lets the novelty wear off, which is the point: performance in week five, when nobody is excited any more, is the best available preview of month six. Past ten weeks, a pilot becomes an unpriced production deployment with no owner, which is the worst of both.

At the end, the written criteria produce one of three outcomes. Convert, with the redesigned process going live and the pilot owner keeping the system. Fix and re-run, with the gap named and a second window. Or stop cleanly, with reasons recorded, a respectable outcome that costs a fraction of a slow fade. What the criteria refuse is the fourth outcome behind both studies' numbers: indefinite extension, in which the pilot is neither failed nor scaled and gradually becomes furniture.

One more MIT finding deserves a place in the scoping conversation: tools built with external partners succeed roughly twice as often as internal builds. The researchers link this to the learning gap, since specialised partners bring systems already designed to adapt to a specific workflow, together with the accumulated pattern-knowledge of previous deployments. Internal teams can build this capability, and some do. The honest scoping question is whether yours already has, and the pilot will tell you cheaply if you let it.

What we might be wrong about

We hold the view that pilot design predicts conversion better than pilot results do, and our engagements keep confirming it. What we are less certain of is the sceptic principle. Including resistant users early has clearly improved the systems we ship; whether it improves conversion, or occasionally kills pilots that would have matured into acceptance, is harder to see from inside individual engagements. Some resistance is information, and some is weather that passes. We are still learning to tell them apart at the scoping stage, and we would rather say so than pretend the distinction is settled.

Mitochondria is an agentic AI product company based in Amsterdam and Pune. ISO 27001:2022 certified. Designed to fall within the limited and minimal risk tiers of the EU AI Act, with controls aligned to the GDPR, UK GDPR and India's DPDP Act.

Next
Next

Why Enterprise AI Underdelivers, and What Fixes It