Most AI pilots don't fail in the demo. The demo goes fine. Somebody feeds a model a stack of contracts, it summarises them in seconds, and the room agrees this changes everything. The pilot fails about three weeks later, quietly, when the person who ran it goes on holiday and nobody else picks it up.
That is not a technology failure. It's a handoff failure, and it's the single most common reason AI work stalls inside mid-sized companies. The model was never the hard part. The hard part is that nobody decided whose job the output becomes.
The Demo Is the Easy Part
A demo has ideal conditions. Clean inputs, someone motivated standing over it, and no consequences if the answer is wrong. Production has none of those. Inputs arrive half-finished from a customer, the person running it has four other things on, and a wrong answer goes out on company letterhead.
The gap between those two states is not closed with a better model. It's closed with the same unglamorous work any other system needs: who runs it, how often, what happens when it's wrong, and who notices if it stops.
Ask a simple question of any pilot you're currently funding. If it produced a wrong answer next Tuesday, who would catch it? If the answer is "nobody, probably," you don't have a pilot. You have a demo that hasn't ended yet.
Nobody Owns the Output
Traditional software arrives with an owner. Someone requested the invoicing system, someone is accountable when invoices are wrong, and everyone knows who to call. AI pilots often arrive sideways — an enthusiastic person, a spare afternoon, a corporate card. The work is real, but it never gets attached to a role.
So the output lands in a grey zone. It isn't quite trusted enough to act on without checking, and checking takes nearly as long as doing the task did. That's the trap: a tool that is genuinely impressive and still a net cost, because the verification burden landed on someone who was never given time for it.
The fix is boring and it works. Name an owner before you name a model. Give them the time, not just the task. And write down what the tool is allowed to decide on its own versus what it only drafts for a human — because "it depends" means every single use is a judgement call, and judgement calls don't scale.
The Data Wasn't the Blocker
We get called in regularly to fix "a data problem" and find a process problem wearing a data problem's coat. The data is messy, certainly. It's always messy. But the pilot didn't die because the data was messy — it died because nobody agreed what the clean version should look like, or who maintains it once it is.
This matters because "we need to fix our data first" is the most expensive way to postpone a decision. It sounds responsible. It can absorb a year. And at the end of it you often have tidier data and exactly the same unanswered question about ownership.
A better sequence: pick one workflow that a named person runs every week. Improve that one with whatever data you have today, mess included. You'll learn what actually needs cleaning within a fortnight, and you'll learn it from use rather than from a workshop.
What a Pilot That Survives Looks Like
The ones that make it past the first quarter tend to share a shape. They're attached to a workflow that already existed, so success is measurable against how it went before. They have one owner who can explain what it does without slides. Failure is visible — someone gets an alert, not a vague sense that quality slipped. And they start narrow enough to be genuinely finished, rather than broad enough to be permanently promising.
That last one catches good teams out. A pilot scoped to "improve customer communications" cannot be completed. A pilot scoped to "draft the first response to warranty emails, for review by the service desk, within one working hour" either works or doesn't, and you'll know inside a month.
Start With the Handoff
If you're planning AI work this quarter, do the handoff design before the model selection. Decide who receives the output, in what form, and what they're expected to do with it. Decide how it fails and who hears about it. Decide when you'd switch it off.
Get that right and the model choice becomes what it should be — a detail you can revisit, not the thing the whole investment rests on. Get it wrong and you'll run another impressive demo next quarter, for a room that's slightly harder to convince.