# Data lake · Glossary
A storage repository that holds large volumes of raw data in its original formats, structured and unstructured, until it is needed.
[Glossary](/glossary) · Data platforms and engineering

# Data lake

     A storage repository that holds large volumes of raw data in its original formats, structured and unstructured, until it is needed. Data lakes are typically built on low-cost object storage.

## Why it matters

     Data lakes make it affordable to retain data that has no immediate use, such as logs, documents, images and sensor data, which is often valuable for machine learning later.

     Without catalogues, ownership and quality controls, lakes become hard to navigate and trust. The value comes from the governance and structure layered on top, not from storage alone.

## In practice

     For example, a UK rail infrastructure contractor might land track-inspection video, sensor readings and maintenance logs in a data lake, with each source catalogued and access-controlled, before selecting subsets for defect-detection models.

## Where Rodan fits

     Rodan designs storage and governance for raw and unstructured data in [Platform and Cloud Engineering](https://rodan.io/platform-cloud-engineering) work.

## Related terms

- [Data warehouse](/glossary/data-warehouse)

- [Data lakehouse](/glossary/data-lakehouse)

- [Unstructured data](/glossary#unstructured-data)

- [Big data](/glossary#big-data)

- [Data catalogue](/glossary/data-catalogue)
HTML: https://rodan.io/glossary/data-lake
