# How to get your data ready for AI in six practical steps
Data ready for AI? Six practical steps to assess quality, governance and pipeline readiness before committing budget to any AI initiative.
Published: 2024-11-20
Author: Rodan Analytics
 Most organisations pursuing AI are not held back by the technology. They are held back by the data underneath it.

 The typical pattern looks like this: a leadership team approves an AI initiative, a vendor gets selected, and then the project stalls. Not because the model is wrong, not because the use case is unclear, but because nobody has a confident answer to the question: "What data do we actually have, and is it good enough to build on?"

 The mistake most organisations at this stage make is treating data readiness as a technical task to hand off after the strategy is agreed. It is not. It is a strategic question that determines whether the strategy is even viable.

 This article will give you a practical, sequenced approach to assessing and improving your data position before committing significant budget to AI. Six steps. No theory. Just what experienced practitioners actually check before recommending that a client moves forward.

---

## Step 1: Define what AI needs to do before you assess what data you have

 Data readiness is not an absolute state. Data that is perfectly adequate for automating invoice processing may be completely inadequate for forecasting customer churn. Before you audit anything, you need to pin down the specific decisions or processes the AI system is intended to improve.

 This sounds obvious. Most teams skip it anyway.

 A manufacturing business we worked with had spent three months preparing a data readiness report before anyone asked which AI application they were actually preparing for. The report assessed data quality across twelve domains. None of the twelve were relevant to the demand forecasting use case the commercial team actually wanted to build.

 Start with the use case. Then work backwards. What inputs does that model or system need? At what frequency? At what level of granularity? That gives you a specific list of data requirements to test against — not a general assessment of data health.

---

## Step 2: Inventory what you own and where it lives

 Most organisations have more data than they think, and less usable data than they hope.

 A practical inventory needs to answer four questions for each potential data source:

- What entity or event does this data describe?

- How is it captured, and who owns it?

- How far back does it go, and at what frequency is it updated?

- What system does it live in, and can that system be accessed programmatically?

 For mid-market businesses, the honest answer to question four is often "it lives in a system that was never designed to be queried." Data is locked in ERP exports, legacy CRM reports or spreadsheets maintained by individuals rather than systems. That is a solvable problem — but you need to know it exists.

 Do not rely on IT to produce this inventory alone. The people who know where the data actually lives are frequently in operations, finance and commercial functions. A cross-functional session of two to three hours will surface more than a month of IT documentation.

---

## Step 3: Assess quality across the dimensions that matter for AI

 General data quality frameworks are too broad to be useful here. For AI specifically, the dimensions that matter most are:

 **Completeness.** Are key fields populated? A customer dataset where 40% of records have no transaction history is not a training set — it is a source of bias.

 **Consistency.** Are the same entities described the same way across systems? A retailer with seven regional ERPs often finds that the same supplier has seven different naming conventions. Joining those records without reconciliation produces garbage.

 **Accuracy.** Does the data reflect reality? This is the hardest to assess because you are testing against ground truth you may not have. Spot checks, triangulation against known outcomes and domain expert review are the practical tools here.

 **Timeliness.** How old is the data, and does its age matter for your use case? A model trained on customer behaviour from 2021 may produce dangerously wrong outputs if buying patterns shifted materially in the intervening period.

 **Volume.** Do you have enough labelled examples for supervised learning, or enough historical observations for a forecasting model? The minimum viable volume varies by model type, but if you have fewer than a few thousand relevant records for a classification task, a bespoke model is probably not where you should start.

 Score each source against each dimension. You are not looking for perfection. You are looking for a clear view of where the gaps are and how much remediation they require.

---

## Step 4: Resolve governance and access before engineering starts

 AI projects fail in production — not in the prototype — because nobody resolved the governance questions during the build phase.

 The questions to answer before you start engineering work:

- Who owns each data source, and do they have authority to permit its use in a model?

- Does the intended use comply with your data processing agreements, privacy policies and any sector-specific regulation (FCA, ICO, GDPR)?

- If the model produces outputs that affect individuals — pricing, credit, hiring — have you assessed the fairness and explainability requirements?

- Where will training data be stored, and who has access to it?

 A private equity-backed retail group we supported discovered mid-project that their most valuable behavioural dataset was governed by a third-party data licence that explicitly prohibited use in model training. That was a six-week delay and a renegotiation nobody had budgeted for.

 Governance review is not a legal formality. It is risk management. Do it at step four, not step ten.

---

## Step 5: Build the pipelines, not just the dataset

 A clean dataset prepared for a proof of concept is not the same thing as a production-ready data pipeline. This distinction is where more AI projects fall down than any other.

 The proof of concept works. The team is confident. The decision is made to scale. And then the question nobody asked during the PoC surfaces: how does fresh data get into this system, cleaned and formatted, on a reliable schedule, without a data engineer manually refreshing a spreadsheet?

 Before you commit to full deployment, you need to know:

- How the data will be ingested from source systems automatically

- What transformation and validation logic will run on each ingestion

- Where that processed data will be stored and in what format

- How failures and anomalies will be detected and handled

- Who is responsible for maintaining the pipeline once it is in production

 This is where tools matter. But the tooling choice should follow the architecture decision, not lead it. Many mid-market organisations do not need a complex modern data stack — they need a well-designed, maintainable pipeline built on infrastructure they already own.

---

## Step 6: Establish a baseline so you can measure what AI actually changes

 This step is almost always omitted. It is also the one that determines whether you can demonstrate return on investment twelve months from now.

 Before your AI system goes live, record the current state of the process it is replacing or augmenting. If the system will improve demand forecasting accuracy, measure your current forecast error. If it will reduce manual review time, time the current review process. If it will improve lead scoring, document your current conversion rate by lead source.

 Without a baseline, you are left with a system that may be performing well and no way to prove it. That is a problem when a new CFO arrives, when the board asks for evidence of returns, or when a budget review puts your AI programme on the table.

 The baseline does not need to be complex. It needs to be specific, measurable and agreed with the stakeholders who will judge the programme's success.

---

## The cost of skipping these steps

 None of these steps are technically difficult. The reason organisations skip them is pace — there is pressure to show progress, to hit a board commitment, to launch before a competitor does. That pressure is real. But a poorly scoped AI initiative built on poor data does not just fail quietly. It produces outputs that erode trust in AI across the organisation, making the next initiative harder to fund and harder to staff.

 The six steps above typically take four to eight weeks for a focused team to work through. That is not a delay. That is the difference between an AI programme that delivers measurable commercial value and one that produces an expensive proof of concept that never reaches production.

 If you want an independent assessment of where your data position actually stands, Rodan offers a focused diagnostic engagement to evaluate your data readiness, identify the highest-value AI use cases for your situation and give you a clear action plan before any significant build budget is committed. Speak to us before you start the build.

---

 **Meta description:** Data ready for AI? Six practical steps to assess quality, governance and pipeline readiness before committing budget to any AI initiative.
HTML: https://rodan.io/insights/how-to-get-your-data-ready-for-ai-in-six-practical-steps
