AI

From Pilots to Production: What Actually Scales

Why most enterprise AI stalls — and the handful of things that move it into production.

May 2026 · 5 min read

Every large organization we talk to has AI pilots. Very few have AI in production. The gap between the two is now the defining feature of enterprise AI — and it is widening, not closing, even as the underlying models get dramatically better.

The conventional explanation is that the technology isn't ready. Our experience says otherwise. The models are ready for far more than most companies ask of them. What stalls is almost never the model. It is everything around the model.

Where pilots go to die

The pattern repeats across industries. A team builds a promising demo in six weeks. Leadership is impressed. Then the pilot enters the space between the demo and the business — and stops moving.

Three walls account for most of the stalls we see.

The data wall. The demo ran on a clean extract. Production has to run on the real thing: inconsistent systems of record, fields that mean different things in different regions, contracts scanned in 2009. Companies discover that their AI project is actually a data project, and data projects have owners, budgets, and politics.

The workflow wall. A model that drafts an answer is a demo. A system that receives the input, checks it, acts on it, logs it, escalates the exceptions, and hands the result to the next step of the process — that is production. The distance between the two is not model quality. It is engineering, process design, and the willingness to change how work actually flows.

The accountability wall. Nobody stops a pilot. Somebody has to sign off on production: who is responsible when the system is wrong, what gets reviewed by a human, what the audit trail shows. Organizations that haven't answered these questions don't say no to AI. They just never quite say yes.

What the successful ones do differently

The organizations that get to production share habits that have very little to do with picking the best model.

They start from a workflow, not a capability. Not "what can AI do?" but "this process costs us 4,000 hours a quarter — take a piece of it end to end." Narrow and deep beats broad and shallow every time.

They put a senior owner on it. Not a committee, not an innovation lab at arm's length from operations — a principal with the authority to change the process the AI is entering. In our engagements, this single factor predicts outcomes better than any technology choice.

They build the boring parts first. Evaluation, logging, exception handling, and rollback are not what demos are made of, but they are what production is made of. Teams that treat reliability as a first-class feature ship; teams that treat it as a cleanup phase don't.

They measure against the process, not the model. Accuracy on a benchmark is a lab number. Cycle time, error rates, and cost per case are business numbers. Production systems are judged — and funded — on the second kind.

The operator's advantage

There is a reason this matters more in energy, trade, and industrial businesses than almost anywhere else. These are domains where the workflows are decades old, the data lives in contracts and counterparty relationships as much as in databases, and the cost of a confident wrong answer is measured in cargoes, not clicks.

That is also why generic AI playbooks underperform here. Moving from pilot to production in these industries requires people who understand both the technology and the physical business it is entering — what a demurrage claim is, why a counterparty's paperwork looks the way it does, which exceptions are routine and which are fires.

The firms that scale AI will not be the ones with the most pilots. They will be the ones that picked fewer, finished them, and let each production system fund the appetite for the next. The technology is ready. The question is whether the organization is.

Start a conversation.

If this touches something you're working on, we'd be glad to talk it through.

Get in touch