Glossary · Agentic systems and generative AI

Token

The unit of text a language model reads and writes, often a word or a fragment of a word. Context window limits, model pricing and response speed are all usually measured in tokens.

Why it matters

Tokens translate directly into cost and latency. Long prompts, large retrieved documents and verbose outputs all consume tokens on every request, which matters when a workflow runs thousands of times a day.

Token counts also vary by language and content type, so estimates based on English prose may not hold for code, tables or other languages. Measuring actual usage in production is more reliable than planning assumptions.

In practice

For example, a UK contact-centre summarisation service might reduce cost by sending only the relevant part of each transcript and asking for a fixed-format summary, rather than passing the entire conversation history with every request.

Where Rodan fits

Rodan monitors token usage, cost and latency as part of operating generative systems built through Applied AI Engineering. See also what is the context window in an LLM and why does it matter.

Related terms