
How to run an AI pilot that actually makes it to production
Most AI pilots fail quietly. Not with a dramatic collapse, but with a slow fade - the project that produced a promising demo, got featured in a board presentation and then stopped being talked about six months later. Nobody cancelled it formally. It just stopped mattering.
This is the dominant pattern at mid-market and enterprise level right now. Organisations are running pilots. Very few are scaling them. The gap between proof-of-concept and production is where AI investment goes to die - and the reason is almost never the technology.
The mistake most organisations at this stage make is treating a pilot as a test of the technology rather than a test of the organisation. The question is not whether the model works. The question is whether your business is structured to absorb it, act on it and sustain it. Those are different questions entirely.
This article sets out how to design and run an AI pilot that is built to scale from day one - what to decide before you start, how to structure the work and what to measure if production is genuinely the destination.
Start with the production question, not the pilot question
Most pilots are scoped backwards. A team identifies a use case, builds a prototype, measures accuracy and then asks: "Should we scale this?" That sequence is the problem.
The right question to ask before the pilot starts is: "What would need to be true for this to run in production?" Work backwards from there.
That means defining, upfront, what production actually looks like. Who owns the output? Which system does it feed into? What human decision does it inform or replace? What does failure look like, and who is accountable when it occurs? What does the commercial return need to be to justify ongoing infrastructure and maintenance cost?
A mid-sized logistics business we worked with wanted to pilot an AI-based demand forecasting model. The pilot was technically successful - the model outperformed their existing approach on every accuracy metric. But it was scoped against a data set maintained by a single analyst who was leaving the business. Nobody had mapped the production data pipeline. The model never deployed. The pilot was not a failure of technology; it was a failure of scoping.
Before you commission a pilot, write a one-page production brief. It does not need to be polished. It needs to answer: what changes in the business when this runs at scale, and are we ready for that?
Choose the use case by deployability, not impressiveness
There is a category of AI use case that generates spectacular demos and almost never reaches production. Anything that requires synthesising unstructured data from ten different sources, involves a decision that a regulator might scrutinise, or depends on a clean data foundation that does not yet exist - these make compelling presentations. They rarely make it to live.
For a first or second AI pilot at a business of your scale, the selection criteria should weight deployability heavily. That means:
- The data required already exists and is accessible
- The decision or output the AI produces connects to a workflow someone actually uses today
- The value is measurable within a 90-day window
- You have an owner - a named individual with the authority and incentive to take it to production
Use case attractiveness and use case deployability are inversely correlated more often than most organisations expect. The CFO who wants to start with a fully autonomous financial close process should probably start with automated variance commentary. The COO who wants to transform the contact centre should probably start with agent-assist on one queue.
Boring deployments compound. Impressive demos do not.
A private equity-backed retail business running on thin margins does not need to pilot a generative AI product recommendation engine on day one. It needs gross margin visibility by SKU, delivered in natural language, so buyers can act faster. That is achievable in weeks. It compounds into something material within a quarter.
Instrument the pilot to prove production value, not pilot value
The metrics that prove a pilot "worked" are rarely the metrics that justify a production budget. Accuracy rates, model performance benchmarks and technical KPIs impress data science teams. They do not move CFOs.
From the first day of the pilot, measure against the commercial outcome the production system is supposed to produce. If the use case is reducing manual review time in a compliance function, measure time saved per case and model it at production volume. If the use case is improving forecast accuracy in a supply chain, measure the inventory cost implication of the accuracy improvement. Attach a number the finance director recognises.
This requires more effort in the design phase. It also means that when you arrive at the investment decision - do we build this into production infrastructure? - you are not asking leadership to make a leap of faith. You are showing them a business case with real numbers, validated by a controlled experiment.
Instrumentation also protects you from false positives. A pilot that shows a 15% improvement in model accuracy but cannot demonstrate commercial impact is a pilot that will struggle to convert. The technology may be working. The use case may be wrong.
Structure the pilot measurement in three layers:
- Technical performance - is the model doing what it is supposed to do?
- Operational adoption - are the people who are supposed to use it actually using it?
- Commercial impact - is the business outcome the pilot was designed to produce moving in the right direction?
All three must be green before you recommend production investment. A pilot that scores well on layer one but poorly on layer two has an adoption problem, not a technology problem. That distinction matters enormously for what you do next.
Build the production path into the pilot governance
Pilots fail to scale for organisational reasons far more often than technical ones. The team that ran the pilot does not have budget authority for production infrastructure. The system the AI output needs to feed into is owned by a different department that was not involved. The vendor relationship is structured as a proof-of-concept engagement with no clear path to a production contract.
These are not surprises. They are predictable failure modes that should be resolved in the pilot governance structure before work begins.
Pilot governance for a business at your scale should include, at minimum: an executive sponsor with production budget authority, a business owner for the target workflow, a representative from IT or engineering who understands the production infrastructure and a finance sign-off on what the commercial threshold for production investment is.
This is not bureaucracy. It is the difference between a pilot that produces a recommendation and a pilot that produces a decision.
At the end of the pilot, the governance structure should be able to answer four questions without debate: Did the pilot achieve its commercial threshold? Is the technology production-ready? Do we have an owner for production operations? What is the cost and timeline to deploy?
If any of those four answers require further investigation at the point the pilot concludes, the pilot was not scoped correctly.
What good looks like at scale
A well-run pilot creates three assets: a production-ready technical build, a validated business case and an internal champion with authority and motivation to see it deployed. Most pilots produce only the first, sometimes. The second and third are what actually move organisations.
The businesses that are building genuine AI capability at this stage are not running more pilots than everyone else. They are running fewer, better-scoped pilots - and converting them at a materially higher rate. Each deployment builds institutional knowledge, data infrastructure and internal confidence. That compounds.
The cost of a failed pilot is not just the direct spend. It is the credibility of the next initiative. It is the senior sponsor who will be harder to engage. It is the team that quietly concludes that AI is not for organisations like theirs.
If your current pilot portfolio looks more like a collection of experiments than a pipeline to production, the problem is almost certainly upstream - in how the work was scoped, not in how it was executed.
Rodan runs structured diagnostic engagements designed specifically to assess your AI pilot approach, identify the use cases most likely to reach production and build the commercial case your board needs to commit. If your pilots are not converting, that is the right place to start.




