The Data Lake
Where raw data pours in and comes out clean enough to train on.
How big?Petabytes. A petabyte is about 500 billion pages of text.
Real: What it looks like.
Keys: arrows rotate · + / − zoom · 0 reset · 1–4 views · S signal · T tour · L labels · Space spin. Models are stylised and built from code: proportions are honest, details are simplified.
What it is
A data lake is huge, cheap storage (mostly hard drives in object-storage clusters) that accepts data in any shape: web pages, images, code, sensor readings, logs.
Teams often organize it in zones. Bronze is raw, exactly as it arrived. Silver is cleaned, de-duplicated and filtered. Gold is curated and ready for a specific use, like training a model.
A catalog keeps track of what exists, where it came from, and who's allowed to use it.
Why AI needs it
Models learn only from the data they're shown. Most of the work in AI goes into what happens here: filtering junk, removing duplicates and private info, and balancing what's left.
Every labeled part
- 1
Sources
Websites, apps, sensors, documents, logs. Data arrives messy.
- 2
Bronze zone
Raw and untouched. Kept so you can always reprocess it.
- 3
Silver zone
Cleaned: duplicates removed, broken files dropped, formats unified.
- 4
Gold zone
Curated for a purpose, like a training set. Quality filtered and documented.
- 5
Object storage
Racks of hard drives. Every file is copied or erasure-coded so a dead drive loses nothing.
- 6
Catalog
The library card index: what's here, where it came from, who may use it.
- 7
Training
Gold data streams out to the GPU cluster.
- 8
Analytics
The same lake also answers business questions.
Try it · concept lab
Bronze → Silver → Gold
Move the raw data through the zones.
- The sky looks blue because air scatters blue light more.
- The sky looks blue because air scatters blue light more.
- BUY CHEAP WATCHES!!! click here click here
- �␀�␀ corrupted file ␀�
- Call Sam at 555-0142 about the order.
- Photosynthesis turns light, water and CO₂ into sugar.
- lol
- A transistor is a switch controlled by voltage.
Real pipelines run steps like these over billions of documents: de-duplicate, drop junk, scrub private info, score quality.
Big idea: What you train on matters as much as the hardware. Most data never makes it to gold.
Try it · concept lab
Break some drives
Tap drives to kill them. Which files survive?
Tap drives to break them.
Each file lives on 3 drives. Any 2 can die.
Big idea: At data-lake scale drives fail every day. Spreading copies, or clever math pieces, means nothing is lost.
Swap it: other ways to do the same job
Data warehouse
Strict tables and fast SQL queries. Great for business numbers, bad for images and raw text.
Lakehouse
A lake with warehouse-style tables on top, aiming for the best of both.
Vector database
Stores embeddings so an AI can look up relevant documents while answering (retrieval).
Go deeper
Where does the data come from?
Most training data comes from the internet: websites, code, pictures and videos, plus books, some bought and some copied without asking. Teams try to remove junk and private details. People disagree about what is fair. Some say anything public should be usable for learning, the way a person learns by reading. Others say it is the creators' work, so they should be asked, credited or paid. Courts are still deciding.
Figures are rough, as of 2026.
Talk about it
- Q1
Cleaning data means throwing out junk and private details. What would you throw out before teaching a robot about your school?
- Q2
AI learns from huge piles of writing and pictures people made. Should people be asked first? Why or why not?
For grown-ups: there are no right answers here. Ask a question, then ask "why do you think that?" The reasons matter more than the answer.
Printable question sheet (PDF)