Glossary · Machine learning and MLOps

Model serving

The infrastructure and software that make a trained model available to applications, typically through an API or batch job, with the performance, scaling and security production use requires. It is where a model becomes part of an operational system.

Why it matters

A model that works in a notebook still has to meet production requirements: response time, throughput, cost, availability, access control and logging. Serving design determines whether the model can be relied on in the workflow it supports.

Serving also enables safe change, through techniques such as running a new model version alongside the old one before switching traffic.

In practice

For example, a UK online marketplace might serve its listing-moderation model through an internal API with defined latency targets, automatic scaling during peak listing hours and logging of every prediction for review.

Where Rodan fits

Rodan builds serving infrastructure for models and AI services through Platform and Cloud Engineering.

Related terms