AI red teaming
Structured adversarial testing in which people deliberately try to make an AI system fail, misbehave or leak information. It looks for harmful outputs, security weaknesses, policy breaches and unsafe tool use before real users or attackers find them.
Why it matters
Standard evaluations measure how a system performs on expected inputs. Red teaming probes the unexpected: manipulative prompts, unusual documents, chained tool calls and misuse by insiders. For systems that take actions or handle sensitive data, those failure modes matter most.
Findings should feed into fixes, guardrails and regression tests, and red teaming should be repeated when the model, prompts, tools or data sources change.
In practice
For example, before launching a customer-service agent that can issue refunds, a UK retailer might run a red-team exercise that attempts prompt injection through order notes, social engineering for refunds outside policy and extraction of other customers’ details, then add controls for each weakness found.
Where Rodan fits
Rodan includes adversarial testing in the evaluation of agentic and generative systems built through Applied AI Engineering.

