# How to evaluate an AI vendor: the questions every business should ask
Evaluating an AI vendor? Learn the questions that reveal real delivery capability, model quality and commercial risk — before you sign anything.
Published: 2024-09-05
Author: Rodan Analytics
 Most organisations approach AI vendor selection the wrong way. They issue an RFP, receive polished decks, sit through demos built on their own data, and then make a decision based on which vendor told the most compelling story.

 That is not evaluation. That is buying a presentation.

 The mistake is structural. Businesses treat AI procurement like software procurement — comparing features, pricing tiers and integration specs. But AI systems do not fail at the feature level. They fail when the underlying approach does not fit the organisation's data maturity, when the vendor cannot support deployment at the pace the business needs, or when the commercial model creates the wrong incentives over time.

 This article will give you a practical framework for evaluating AI vendors properly. Not a checklist of questions to paste into your next RFP, but a way of thinking about what you are actually buying, where the real risks sit and how to structure conversations that reveal the truth rather than the pitch.

---

## Understand what you are actually buying

 Before you evaluate a vendor, you need to be precise about what category of AI system you are procuring. This sounds obvious. It rarely is.

 There is a meaningful difference between a vendor selling you access to an AI platform, a vendor building a custom AI solution on top of existing infrastructure, and a vendor deploying pre-built AI products configured to your context. Each carries different risk profiles, different cost structures and different dependencies.

 A logistics company evaluating route optimisation tools, for example, might receive proposals from three vendors that all describe themselves as "AI-powered." One is licensing a foundation model API with a thin wrapper. One has built proprietary ML models trained on industry-specific freight data. One is offering a rules-based system with some probabilistic elements and calling it AI. The right questions — and the right answers — are completely different across those three.

 Before any vendor conversation, define:

- What specific business decision or process this system will affect

- What data you already have and what quality it is in

- Whether you need explanation and auditability of outputs, or just accuracy

- Who owns the system over time — vendor dependency versus internal capability

 If you cannot answer these four questions before the process starts, you are not ready to evaluate. You are ready to be sold to.

---

## The questions that reveal delivery capability

 A vendor's ability to build something impressive in a sandbox tells you almost nothing about their ability to deliver in your environment. The most important questions are about deployment, not demonstration.

 Ask vendors to describe, in specific terms, the last three implementations they completed with organisations of similar size and complexity to yours. Not case studies with logos and percentage uplifts — actual accounts of what was hard, what took longer than expected and what changed between the initial proposal and go-live.

 Vendors who have genuinely delivered at scale will answer this without hesitation. Vendors whose experience sits primarily in proof-of-concept work will struggle to give you specifics.

 Push further:

- **Who owns the integration work?** If your systems require custom connectors, data pipeline builds or model fine-tuning on proprietary data, understand exactly who does that work, how it is priced and what happens when it hits problems.

- **What does your team need to provide?** The hidden cost in AI deployments is rarely the vendor's day rate. It is the internal time required — data engineering, subject matter expertise, IT security review, change management. Ask for a realistic estimate of your internal resource requirement.

- **What is the escalation path when something breaks?** This is not a pessimistic question. It is the question that distinguishes vendors who have operated in production environments from those who have not.

 A mid-market financial services firm recently contracted an AI vendor on the strength of a compelling demo using synthetic data. Eight months later, the system was still not in production. The vendor's implementation methodology assumed a clean, structured data environment. The firm's actual data estate — spread across three legacy systems and a data warehouse built in 2014 — bore no resemblance to that assumption. The cost of that mismatch was not just the vendor contract. It was the internal resource absorbed, the opportunity cost of the delayed use case and the political cost of a failed initiative.

---

## How to assess model quality and risk

 Vendor claims about accuracy, performance and reliability are rarely verifiable from the outside. But you can probe the conditions under which their models perform — and the conditions under which they do not.

 Ask three things specifically:

 **First, how was the model evaluated?** Any serious vendor should be able to tell you the evaluation methodology, the benchmark datasets used and the conditions of the test. If accuracy claims come without this context, they are marketing figures.

 **Second, how does the model behave at the edges?** Good AI systems are designed to fail gracefully. Ask what happens when the model encounters data that falls outside its training distribution. Does it return a low-confidence flag? Does it default to a rule-based fallback? Does it hallucinate an answer with no indication of uncertainty? The answer tells you a great deal about the engineering maturity behind the product.

 **Third, what is the retraining cycle?** AI models degrade over time as the real-world distribution of data they were trained on shifts. A vendor who cannot articulate a clear position on model monitoring, drift detection and retraining cadence is selling you a static system in a dynamic environment.

 For organisations in regulated sectors — financial services, healthcare, insurance — add questions about explainability. A model that cannot produce a human-readable justification for its outputs may be commercially useless regardless of its accuracy, because your compliance and governance requirements will not permit it in production.

---

## Commercial model and long-term dependency

 The most dangerous moment in AI vendor selection is when you let a technically impressive proposal override a poorly structured commercial agreement.

 Several commercial models are common in the market. Some vendors charge on a per-seat or per-user basis, which scales poorly as adoption grows. Some charge on consumption — per query, per inference, per document processed — which creates unpredictable cost as usage increases. Some offer outcome-based pricing, which sounds attractive but is rarely structured in a way that genuinely aligns incentives.

 Understand the following before signing:

- **Data ownership.** Who owns the training data, the fine-tuned models and the outputs? Ensure your contract is explicit. Some vendors retain rights to use your data to improve their general models. That may be unacceptable depending on the sensitivity of your data.

- **Portability.** If you want to exit the relationship in two years, can you take your models, your data and your workflow logic with you? What does migration actually cost?

- **Price escalation.** AI infrastructure costs are not stable. Understand the vendor's pricing protections and what triggers a renegotiation.

 One of the most common mistakes at this stage is optimising for the initial contract value rather than the total cost of the relationship. A low headline price with high switching costs is a worse deal than a higher headline price with genuine portability.

---

## Build the evaluation process to surface the truth

 A structured evaluation process — rather than a beauty parade of demos — is the only way to generate comparable, honest signals across vendors.

 Run a structured pilot. Give each shortlisted vendor the same constrained problem, using a subset of your real data, with a defined success criterion and a time limit. This does not need to be expensive. A well-scoped pilot can be completed in two to four weeks and costed under £20,000, including internal time.

 Score vendors on three dimensions: technical performance against the defined criterion, quality of the working relationship during the pilot, and the honesty of their communication when they hit problems. The third dimension is the most important. A vendor who surfaces problems early and proposes solutions is demonstrating the operating culture you will be working with for years.

 Before final selection, take a reference from a client who left the vendor. Most vendor-supplied references are warm. Find out what the experience looks like from a client who made a different decision.

---

## The cost of getting this wrong is not the contract value

 Organisations that select AI vendors carelessly do not just waste money on the wrong contract. They exhaust internal appetite for AI investment, damage their credibility with boards and executive teams, and delay the use cases that would have generated real commercial value.

 For a business between £500m and £1bn in revenue, a failed AI initiative typically costs two to four times its headline contract value when you account for internal resource, opportunity cost and the time required to restart. That is not a reason to avoid AI investment. It is a reason to be rigorous about how you evaluate who you trust to deliver it.

 If you are in the process of evaluating AI vendors now, the right first step is not to extend your RFP. It is to pressure-test your own requirements before you test the vendors.

 Rodan's diagnostic engagement is structured to do exactly that — clarify the use case, assess your data readiness, and define the evaluation criteria that will actually predict delivery success. It costs between £1,000 and £2,000 and typically takes two weeks. [Book a diagnostic with Rodan.](https://rodan.io)

---

 **Meta description:** Evaluating an AI vendor? Learn the questions that reveal real delivery capability, model quality and commercial risk — before you sign anything.
HTML: https://rodan.io/insights/how-to-evaluate-an-ai-vendor-the-questions-every-business-should-ask
