Glossary · Agentic systems and generative AI

Golden dataset

A curated set of representative inputs with agreed correct or acceptable outputs, used as a fixed benchmark to evaluate an AI system. It is the reference against which changes to models, prompts or retrieval are tested.

Why it matters

Without a stable benchmark, teams cannot tell whether a change made things better or worse, and they tend to judge quality by a handful of memorable examples. A golden dataset turns that judgement into repeatable measurement.

It must reflect real usage, including difficult and edge cases, and be reviewed by people who understand the domain. It should be refreshed as the system’s use changes, while keeping earlier versions for comparison.

In practice

For example, a UK conveyancing firm building a lease-review assistant might assemble a few hundred lease clauses with solicitor-approved extractions, including unusual wording, and require every new release to match or improve on the previous score.

Where Rodan fits

Rodan creates evaluation datasets with domain experts at the start of AI and Decision Systems work, so quality is measured from the first iteration.

Related terms