What marketing data do you need before AI adds any value?

What marketing data do you need before AI adds any value?

Most marketing teams buying AI tools are doing it in the wrong order. They are investing in models before they have anything worth modelling. The result is expensive, slow and quietly embarrassing - a lot of dashboards that do not drive decisions, a lot of vendor promises that did not survive contact with the actual data.

The mistake is understandable. AI moves fast and boards are asking questions. But the organisations that get real commercial returns from AI in marketing are not the ones that moved first. They are the ones that built the right data substrate before they applied the intelligence layer.

This article is not about which AI tools to buy. It is about what your data needs to look like before any of those tools can do useful work. If you are a marketing director, head of ecommerce or growth leader at a consumer business, this is the diagnostic you should run before your next vendor conversation.

The data your AI actually needs to function

AI in marketing is not magic. It is pattern recognition applied to historical data. That means the quality of your output is entirely constrained by the quality, completeness and structure of your input.

Before any AI system can generate useful insight - let alone act autonomously - you need four categories of data in reasonable shape.

Customer identity data. Can you stitch a single customer across channels? If a person buys on your website, opens an email and then converts via paid social retargeting, do those three events resolve to one profile? Most mid-market consumer businesses cannot answer yes to this cleanly. They have a CRM, a CDP or a data warehouse that partially solves the problem, but with gap years, duplicate records and missing linkage keys.

Behavioural data. Clickstream, session data, on-site engagement, app events. This data is often collected but rarely structured in a way that supports modelling. Raw GA4 exports are not a behavioural dataset. They are a log file. There is a significant transformation step between collection and readiness.

Transactional data. Purchase history, basket composition, returns, refunds, channel of acquisition. This sounds like table stakes, but the problem is usually fragmentation. Data sits in an ERP, a Shopify instance, a bespoke OMS and a loyalty platform that were never designed to talk to each other.

Attribution data. Not perfect attribution - that does not exist - but a consistent, agreed methodology for how you credit channels and campaigns. If your attribution model changes every quarter because someone installed a new tool, your historical data is not a reliable training set for anything.

None of this requires perfection. But it requires enough consistency and completeness that a model can learn something real from it.

Where the gaps usually sit

Take a direct-to-consumer apparel brand with around £80m in revenue. They had Klaviyo, GA4, Shopify and a third-party attribution tool. On paper, they had the data stack. In practice, their customer identity resolution was running at roughly 40% match rate across channels. Their GA4 implementation had broken event tracking across two major site updates. Their Shopify data had three years of clean transactional history and then a platform migration that had corrupted eighteen months of basket-level data.

When they tried to deploy an AI-driven personalisation engine, the vendor onboarded them, ran the integration and came back six weeks later to say the model could not train reliably. Too much noise, too many gaps, too little signal.

This is not unusual. It is, arguably, the default state for a consumer business at this scale. The fix is not glamorous - it is data auditing, identity resolution work and pipeline remediation. But it is the prerequisite.

The gaps that tend to matter most are:

  1. Identity resolution below 60% match rate across channels
  2. Incomplete or inconsistent event taxonomy across web, app and email
  3. Attribution methodology that has changed more than once in two years
  4. Transactional history under eighteen months at product or basket level
  5. No clean mapping between customer segment and acquisition channel

If three or more of those apply to your business, you are not ready for AI. You are ready for data infrastructure work.

What "good enough" actually looks like

You do not need a perfect data warehouse before you start. You need a defined minimum viable dataset for the specific AI application you are trying to run.

If the goal is churn prediction, you need at minimum: twelve months of transaction history per customer, a stable customer ID, and a signal for what "churned" means in your business (lapsed definition, subscription cancellation, whatever is relevant). That is a narrow, achievable data requirement.

If the goal is next-best-offer personalisation, you need clean product catalogue data, basket-level transaction history and some form of behavioural segmentation. Still achievable, but it requires more groundwork.

If the goal is autonomous media buying - letting an AI agent manage budget allocation across paid channels in real time - you need clean, consistent, near-real-time conversion data flowing from every channel into a single place. That is a harder infrastructure problem and most businesses at the £100m–£500m revenue range are not there yet.

The discipline here is to scope your AI ambition to match your data maturity, then use early wins to build the infrastructure that enables the next application. This is not caution for its own sake. It is sequencing that actually delivers results rather than burning budget on a capability you cannot use yet.

The organisational data problems that AI exposes

Here is the thing that most AI vendor conversations skip: data quality is not a technical problem. It is an organisational one.

The reason identity resolution is broken at your business is usually because the ecommerce team, the CRM team and the paid media team are measured on different KPIs and have never had a commercial reason to agree on a shared customer ID. The reason your attribution methodology changes every quarter is because the CFO and the CMO are having a proxy argument about channel efficiency, and each new tool is a ceasefire.

AI does not fix these problems. It exposes them faster and at greater cost.

Before you invest in AI tooling, ask yourself whether your marketing organisation has a single agreed definition of a customer, a consistent view of lifetime value and an attribution methodology that the CFO and CMO will both defend. If not, those are the conversations to have first.

A senior marketing data leader - whether a full-time hire or a fractional CDO engagement - can drive these decisions in a way that a vendor implementation never will. The tools come after the alignment.

Building the data foundation in practice

The practical sequence looks like this.

First, audit what you have. Not a vendor audit - an independent one. Map every data source, understand match rates, identify the gaps that would break a model.

Second, prioritise by AI use case. Decide which AI application would generate the most commercial value if it worked. Work backwards from that to the minimum data requirements. Do not boil the ocean.

Third, remediate the critical gaps. This usually means identity resolution work, event taxonomy standardisation and - often - a data warehouse consolidation project that was overdue anyway.

Fourth, run a narrow proof of concept. One model, one use case, real data, real commercial outcome as the success metric. Not a demo. Not a pilot with vendor data. Your data, your customers, your revenue.

Fifth, scale what worked. Once you have a proof of concept that holds up under scrutiny, you have the template for the next application.

Rodan's diagnostic engagement is designed precisely for steps one and two. In a structured £1–2k engagement, we will tell you what your data actually looks like, where the gaps sit relative to your AI ambitions and what the remediation roadmap should be - before you spend material budget on tooling.

The cost of skipping this work

The marketing leaders who skip the data foundation step do not usually fail publicly. What happens is quieter: the AI tool gets deployed, produces mediocre results, generates internal scepticism and gets shelved. The business concludes that AI does not work for them. The real conclusion should be that the data was not ready.

That scepticism has a cost. Competitors who did the groundwork are now running AI-driven personalisation, dynamic pricing and predictive retention at lower cost and higher precision. The gap compounds.

Do the diagnostic first. Understand what you have. Fix what matters. Then apply the intelligence layer to something that can actually learn from your data.

If you want an independent assessment of your marketing data maturity before your next AI investment, book a diagnostic with Rodan.