
How to evaluate your current data infrastructure before investing in AI
Most organisations at this revenue scale have already spent serious money on data. A data warehouse here, a BI platform there, maybe a data lake that nobody fully trusts. The question is not whether you have data infrastructure - you do. The question is whether it is fit for what you are about to ask of it.
The mistake most firms make at this stage is treating AI investment as a standalone decision. They evaluate models, vendors and use cases without first asking whether the foundation beneath those use cases will hold. The result is predictable: pilots that work in controlled conditions, then collapse when they touch real production data. Six-figure investments that deliver nothing because the underlying data is incomplete, inconsistent or simply not accessible to the systems trying to use it.
Before you commit budget to AI, you need an honest assessment of what your data infrastructure can actually support. This article gives you a structured way to do that - covering data quality, architecture, governance and organisational readiness - so you can invest in the right sequence.
Why AI amplifies data problems rather than hiding them
There is a persistent belief that modern AI tools are forgiving about data quality. They are not. Traditional BI can mask a bad data environment behind manual processes and analyst workarounds. AI cannot. A large language model querying your commercial data will surface every inconsistency. An agentic system automating a pricing or procurement workflow will propagate errors at scale and at speed.
Consider a distribution business with £800m in revenue running a demand forecasting pilot. Their ERP holds five years of sales history, but three acquisitions over that period mean product codes are not consistently mapped across systems. The forecasting model trains on the merged dataset without that context, produces outputs that look plausible, and feeds into replenishment decisions. Two quarters later, the business is carrying excess stock in categories the model treated as high-velocity when they were actually declining legacy lines absorbed through acquisition.
The problem was not the model. The problem was that nobody had audited what the training data actually represented before the project started.
The implication for senior leaders is straightforward: your AI investment thesis needs a data infrastructure assessment as a precondition, not an afterthought.
The four dimensions of a data infrastructure assessment
A credible pre-AI audit covers four areas. Each one has direct implications for which AI use cases are viable and in what timeframe.
1. Data quality and completeness
Start with the data that will feed your highest-priority AI use cases. For most firms in this revenue range that means commercial data (revenue, margin, customer), operational data (supply chain, fulfilment, production) and financial data (actuals, forecasts, planning).
For each domain, assess:
- Completeness: what percentage of records are fully populated across the fields your use cases require?
- Consistency: where the same entity (a customer, a product, a supplier) appears in multiple systems, does it carry consistent identifiers and attributes?
- Timeliness: how fresh is the data when it arrives in your analytical environment, and does that latency matter for the use case?
- Lineage: can you trace a figure back to its source, and does anyone in the organisation actually own the definitions?
A private equity-backed consumer brand discovering that its customer lifetime value calculation uses three different definitions across finance, marketing and the data team - each producing materially different numbers - is not ready to build an AI-driven retention model. It is ready to fix its definitions first.
2. Architecture and accessibility
The second dimension is whether your data can physically reach the systems that need to use it, in the form they need it, at the latency they require.
The questions to ask are:
- Where does your data live - cloud, on-premise or hybrid - and what are the movement costs and constraints?
- Do you have a functioning semantic layer, or does every analyst query raw tables with their own logic?
- What are your integration patterns between source systems and your analytical environment, and are they reliable enough for automated pipelines?
- Do you have the compute resources to support model inference at the frequency your use cases demand?
A business that wants to deploy a real-time commercial intelligence tool - the kind of capability that lets a commercial director ask natural-language questions of live revenue data - needs sub-minute data freshness and a governed semantic layer. If your current architecture refreshes overnight and your semantic definitions live in a shared spreadsheet, you have an architecture gap, not an AI gap.
3. Governance and trust
Data governance is the area most commonly underestimated at this revenue level, because it looks administrative until it becomes a liability.
Before deploying AI, you need clear answers to: who owns each data domain, who has the authority to change definitions, and how access is controlled and audited. Regulatory exposure matters here - financial services firms, healthcare businesses and any organisation processing personal data at scale need governance that can demonstrate compliance to an AI system's outputs, not just its inputs.
The practical risk in under-governed environments is that AI outputs become authoritative faster than the organisation's ability to challenge them. A board that starts making capital allocation decisions based on AI-generated scenario outputs needs to know those outputs rest on definitions and data that have been formally validated. Without governance, they cannot know that.
4. Organisational and skills readiness
Infrastructure is not only technical. The organisation consuming the outputs is part of the system.
Assess whether your data team has the capability to maintain AI pipelines, retrain models when data distributions shift and monitor for output degradation. Assess whether business users have enough data literacy to interrogate AI outputs rather than accept them. Assess whether your decision-making processes are structured to act on automated insight at the speed AI can generate it.
A retailer with strong data engineering capability but no ML operations practice will struggle to keep a deployed model performing over time. A business with sophisticated models but a leadership team that does not trust the outputs will never realise the value of the investment.
How to prioritise remediation before committing to AI
Once you have assessed all four dimensions, you will rarely find a clean bill of health. The goal is not perfection - it is fitness for purpose relative to specific use cases.
Use a simple two-axis framework: rate each AI use case by business value (revenue impact, cost reduction or risk reduction) against data readiness (the proportion of the four dimensions that are already in good shape for that use case). High-value use cases with high data readiness are your near-term plays. High-value use cases with low data readiness define your remediation roadmap.
This framing matters for commercial conversations. When a business case lands on the CFO's desk for an AI investment, the right challenge is not "why are we doing this?" - it is "what does data readiness look like for this use case, and what is the remediation cost if it is not there yet?"
Folding remediation costs into the AI business case upfront, rather than discovering them mid-delivery, changes both the investment decision and the timeline significantly.
What good looks like at this revenue scale
Firms in the £500m to £1.5bn range that are well-positioned for AI investment typically share a small number of characteristics. They have a single, trusted source of truth for their most critical commercial and operational metrics - usually a cloud data warehouse with a maintained semantic layer. They have data ownership assigned at the domain level, with definitions documented and change-controlled. They have at least one team with practical experience of building and maintaining data pipelines. And they have leadership that treats data quality as an operational discipline rather than a technology project.
None of that is exotic. Most of the firms Rodan works with are partially there - they have made the investments, but execution has been inconsistent. The assessment process makes that visible, so remediation can be targeted rather than speculative.
The cost of skipping this step
The firms that skip infrastructure assessment before AI investment do not usually fail visibly. They succeed partially - a pilot that works, a proof of concept that impresses - and then plateau. They spend the next eighteen months troubleshooting problems that a two-week diagnostic would have surfaced before a penny was committed to models or platforms.
For a business at £1bn in revenue, an AI initiative that stalls at proof-of-concept and never reaches production scale is not just a wasted project cost. It is a leadership credibility problem, a missed competitive window and often a setback for data investment more broadly.
The sequence matters: assess, remediate, deploy. Not the other way around.
If you want a structured view of where your data infrastructure stands before making your next AI investment, Rodan's diagnostic engagement is built precisely for this - a fixed-scope, fixed-cost assessment that tells you what is ready, what is not and what to do about it.




