Skip to content
Hi, Bot

The Data Lake

Where raw data pours in and comes out clean enough to train on.

How big?Petabytes. A petabyte is about 500 billion pages of text.

Real: What it looks like.

Keys: arrows rotate · + / − zoom · 0 reset · 1–4 views · S signal · T tour · L labels · Space spin. Models are stylised and built from code: proportions are honest, details are simplified.

What it is

A data lake is huge, cheap storage (mostly hard drives in object-storage clusters) that accepts data in any shape: web pages, images, code, sensor readings, logs.

Teams often organize it in zones. Bronze is raw, exactly as it arrived. Silver is cleaned, de-duplicated and filtered. Gold is curated and ready for a specific use, like training a model.

A catalog keeps track of what exists, where it came from, and who's allowed to use it.

Why AI needs it

Models learn only from the data they're shown. Most of the work in AI goes into what happens here: filtering junk, removing duplicates and private info, and balancing what's left.

Every labeled part

  1. 1

    Sources

    Websites, apps, sensors, documents, logs. Data arrives messy.

  2. 2

    Bronze zone

    Raw and untouched. Kept so you can always reprocess it.

  3. 3

    Silver zone

    Cleaned: duplicates removed, broken files dropped, formats unified.

  4. 4

    Gold zone

    Curated for a purpose, like a training set. Quality filtered and documented.

  5. 5

    Object storage

    Racks of hard drives. Every file is copied or erasure-coded so a dead drive loses nothing.

  6. 6

    Catalog

    The library card index: what's here, where it came from, who may use it.

  7. 7

    Training

    Gold data streams out to the GPU cluster.

  8. 8

    Analytics

    The same lake also answers business questions.

Try it · concept lab

Bronze → Silver → Gold

Move the raw data through the zones.

8/8 rows kept
  • The sky looks blue because air scatters blue light more.
  • The sky looks blue because air scatters blue light more.
  • BUY CHEAP WATCHES!!! click here click here
  • �␀�␀ corrupted file ␀�
  • Call Sam at 555-0142 about the order.
  • Photosynthesis turns light, water and CO₂ into sugar.
  • lol
  • A transistor is a switch controlled by voltage.

Real pipelines run steps like these over billions of documents: de-duplicate, drop junk, scrub private info, score quality.

Big idea: What you train on matters as much as the hardware. Most data never makes it to gold.

Try it · concept lab

Break some drives

Tap drives to kill them. Which files survive?

Tap drives to break them.

File A: readableFile B: readableFile C: readableFile D: readable

Each file lives on 3 drives. Any 2 can die.

Big idea: At data-lake scale drives fail every day. Spreading copies, or clever math pieces, means nothing is lost.

Swap it: other ways to do the same job

  • Data warehouse

    Strict tables and fast SQL queries. Great for business numbers, bad for images and raw text.

  • Lakehouse

    A lake with warehouse-style tables on top, aiming for the best of both.

  • Vector database

    Stores embeddings so an AI can look up relevant documents while answering (retrieval).

Go deeper

Where does the data come from?

Most training data comes from the internet: websites, code, pictures and videos, plus books, some bought and some copied without asking. Teams try to remove junk and private details. People disagree about what is fair. Some say anything public should be usable for learning, the way a person learns by reading. Others say it is the creators' work, so they should be asked, credited or paid. Courts are still deciding.

Figures are rough, as of 2026.

Talk about it

  1. Q1

    Cleaning data means throwing out junk and private details. What would you throw out before teaching a robot about your school?

  2. Q2

    AI learns from huge piles of writing and pictures people made. Should people be asked first? Why or why not?

For grown-ups: there are no right answers here. Ask a question, then ask "why do you think that?" The reasons matter more than the answer.

Printable question sheet (PDF)