# Why your AI project needs a data strategy first
Why AI projects fail before they start — and why data strategy must come before AI investment for firms between £500m and £1.5bn revenue.
Published: 2025-09-22
Author: Rodan Analytics
 Most AI projects fail before a single model is trained. Not because the technology is wrong, not because the team lacks capability, but because the organisation trying to run the project has no coherent view of its own data.

 The pattern is familiar. A senior leader approves a budget. A vendor is engaged or an internal team assembled. Ambitions are set around automation, forecasting or customer intelligence. Then, three months in, the project stalls. The data is incomplete. It lives in four systems that do not talk to each other. Nobody agreed on definitions. The finance team's "revenue" is not the same as the commercial team's "revenue." The AI cannot help you if it cannot trust what it is reading.

 The mistake most organisations at this stage make is treating data infrastructure as a downstream concern — something to sort out once the AI use case is proven. It is actually the precondition.

 This article sets out why data strategy must come before AI investment, what that strategy needs to cover and how to sequence the work practically.

---

## The problem with starting with the use case

 The instinct to begin with a use case is understandable. Use cases are concrete, they can be presented to a board, and they give the project a commercial frame. But they create a trap.

 When you define the AI use case first, you anchor the entire programme to that use case. The data work that follows becomes scoped to serve it. If the use case changes — and it will — the data work has to restart. More importantly, use-case-first thinking rarely surfaces the deeper data quality and governance problems that will block execution. Those only emerge when someone tries to actually build something.

 Consider a distribution business in the £700m revenue range attempting to build a demand forecasting model. The use case is well defined and commercially important. But when the data science team begins the build, they discover that product codes changed twice in three years and nobody reconciled the historical records. Promotional data lives in spreadsheets held by individual account managers. Customer return rates — a key input — are tracked differently across three warehouse management systems. The forecasting model cannot be trained reliably until someone fixes those problems. That work takes longer than building the model itself.

 This is not an unusual scenario. It is the norm. The use case was right. The sequence was wrong.

---

## What a data strategy actually is

 A data strategy is not a technology architecture document. It is not a list of tools to buy. It is a set of decisions about what data your organisation needs to run commercially, how that data will be owned, governed and maintained, and what infrastructure is required to make it accessible and trustworthy.

 For a firm between £500m and £1.5bn revenue, a practical data strategy covers five things:

- **Data domains and ownership** — which business functions own which data, and who is accountable for quality within each domain

- **Definitions and standards** — agreed canonical definitions for the metrics that matter commercially (revenue, margin, customer, churn, etc.)

- **Data flows and integration** — how data moves between systems, where duplication exists and where the authoritative source of record sits

- **Quality and governance** — how data quality is measured, what the acceptable thresholds are and what the remediation process looks like

- **Access and enablement** — how business users and analytical systems can query data reliably, without dependence on a small number of technical gatekeepers

 None of this requires perfection before AI work begins. It requires enough clarity that the AI system has a stable, trustworthy foundation to operate on.

 A useful test: if you cannot answer "what is our single source of truth for customer revenue, and who owns it?", you are not ready to deploy an AI system that makes commercial decisions using customer revenue data.

---

## The governance question nobody wants to have

 Data governance is the part of the conversation that tends to kill momentum in the room. It sounds bureaucratic. It feels like a distraction from the interesting work. Senior leaders nod and delegate it.

 That is the wrong response. Governance is not administrative overhead. It is the mechanism by which AI outputs become trustworthy enough to act on.

 An AI system that produces a customer churn prediction is only useful if the people receiving that prediction trust the underlying data. If the sales team knows that the CRM data is unreliable — because they are the ones who enter it and they know how inconsistently it gets done — they will not change their behaviour based on the model's output. The model becomes an expensive report that nobody reads.

 Governance also becomes a legal and commercial issue at scale. Firms in regulated sectors — financial services, healthcare, energy — face real exposure when AI-driven decisions cannot be traced back to auditable, well-defined data sources. The question "where did this recommendation come from, and can you defend it?" is one that regulators and boards will increasingly ask.

 The governance work does not need to be a two-year programme before anything moves. A focused engagement — typically four to eight weeks — can establish the ownership model, the critical definitions and the quality baselines needed to proceed with AI development safely. Starting there is not a delay. It is risk management.

---

## How to sequence the work

 Getting the sequence right matters as much as getting the strategy right. Here is a practical approach for organisations at this stage.

 **Start with a data audit, not a data warehouse.** Understand what you have before you decide what to build. Map the critical data domains, identify the systems of record, surface the quality gaps and duplication. This takes weeks, not months, and it shapes every decision that follows.

 **Define commercial outcomes, not technical requirements.** The data strategy should be driven by the decisions the business needs to make, not the capabilities of the tools available. If the CFO needs better visibility on margin by customer segment, start there. What data underlies that? Where does it live? Is it trustworthy?

 **Prioritise the data that feeds your highest-value AI use cases.** You do not need to fix everything. You need to fix the data that sits under the two or three AI applications that will generate the most commercial value in the next 12 months. Sequence data remediation around commercial priority.

 **Build for reuse, not for the project.** The data infrastructure you build to support your first AI use case should be designed to support the next five. A well-structured data layer — clean, governed, accessible — compounds in value. Every subsequent AI project becomes faster and cheaper to deliver.

 **Establish baseline reporting before you build predictive models.** If your business cannot reliably report on what happened last month, it cannot reliably predict what will happen next month. Organisations that skip this step discover the problem when their AI outputs contradict their management accounts and nobody knows which number to trust.

---

## The cost of getting this wrong

 Organisations that skip data strategy and go straight to AI implementation do not save time. They lose it. The average mid-market AI project that stalls on data issues loses six to twelve months before leadership accepts that the foundational work needs to happen. By then, the budget has been partially consumed, team morale has taken a hit and the business case has to be rebuilt.

 There is also a compounding cost. Poor data infrastructure does not just slow AI projects — it degrades every analytical and operational process that depends on it. The finance team's month-end close, the commercial team's forecasting, the ops team's capacity planning: all of these run worse when data ownership is unclear and quality is inconsistent.

 The firms that get the most from AI investment are rarely the ones with the most sophisticated models. They are the ones with the cleanest data, the clearest ownership and the governance discipline to maintain both. That is a strategy problem before it is a technology problem.

 If you are planning an AI programme in the next six to twelve months and have not yet done the foundational data work, the right first step is a structured diagnostic. Rodan's data and AI readiness diagnostic — a focused, fixed-fee engagement — maps your current data maturity, identifies the gaps most likely to block execution and gives you a clear sequencing plan before any material investment is made. [Book a diagnostic with Rodan.](https://rodan.io)

---

 **Meta description:** Why AI projects fail before they start — and why data strategy must come before AI investment for firms between £500m and £1.5bn revenue.
HTML: https://rodan.io/insights/why-your-ai-project-needs-a-data-strategy-first
