# Data labelling · Glossary
Adding the correct answer or category to raw examples, such as tagging an email as a complaint or marking defects in an image, so they can be used to train or evaluate a model.
[Glossary](/glossary) · Machine learning and MLOps

# Data labelling

     Adding the correct answer or category to raw examples, such as tagging an email as a complaint or marking defects in an image, so they can be used to train or evaluate a model. Labelling can be done by experts, trained annotators, existing process outcomes or AI-assisted tools.

## Why it matters

     Labels encode the organisation’s definitions. If annotators interpret categories differently, the model learns that inconsistency. Clear guidelines, agreement checks between labellers and review of disagreements matter more than volume.

     Labelling is also a governance activity: it may expose annotators to personal or sensitive data, and the resulting labels become an asset that needs versioning and ownership.

## In practice

     For example, a UK energy supplier building a complaint classifier might have experienced agents label a sample of customer emails against a short, written taxonomy, measure how often two agents agree, and refine unclear categories before labelling more.

## Where Rodan fits

     Rodan designs labelling guidelines and review workflows with subject-matter experts in [AI and Decision Systems](https://rodan.io/what-we-build/ai-decision-systems) work.

## Related terms

- [Ground truth](/glossary/ground-truth)

- [Training data](/glossary/training-data)

- [Supervised learning](/glossary/supervised-learning)

- [Golden dataset](/glossary/golden-dataset)

- [Human-in-the-loop (HITL)](/glossary#human-in-the-loop)
HTML: https://rodan.io/glossary/data-labelling
