A Training Run
Thousands of GPUs learning one model together for weeks.
How big?Weeks to months, thousands to 100,000+ GPUs.
Real: What it looks like.
Keys: arrows rotate · + / − zoom · 0 reset · 1–4 views · S signal · T tour · L labels · Space spin. Models are stylised and built from code: proportions are honest, details are simplified.
What it is
Data streams from the lake through CPU data loaders into the GPUs in batches. Each GPU guesses, measures how wrong it was, and works out how to nudge the model's weights to be less wrong.
Then comes the all-reduce: every GPU shares its nudges over the network and they're averaged, so all copies of the model stay identical. This happens for every step, millions of times.
Every few hours the run saves a checkpoint. If a GPU dies (and at this scale one always does), training restarts from the last save.
Why AI needs it
This is how a model like a chatbot gets made. The hardware's job is to keep every GPU busy and never waiting on data, memory or the network.
Every labeled part
- 1
Data lake
Gold-zone training data.
- 2
Data loaders
CPUs read, shuffle and batch the data so GPUs never starve.
- 3
GPU servers
Each runs a copy (or a slice) of the model on its batch.
- 4
Network fabric
Leaf and spine switches. Explode to lift them above the nodes.
- 5
All-reduce ring
Light-blue pulses: every GPU passing its share of the update around the ring until all agree.
- 6
The model
Billions of weights, slowly getting better at predicting.
- 7
Checkpoints
Saved progress, so a failure costs hours, not weeks.
Try it · concept lab
Agreeing on an answer: the ring
Pass numbers around until every GPU knows the total.
Passes
0 / 10
Each GPU only ever talks to its neighbour, so no single link gets swamped. Real all-reduce splits the numbers into chunks and does both phases on every link at once; the idea is the same.
Big idea: After every training step, all GPUs must agree. Passing around a ring keeps every link equally busy, so it scales to thousands of GPUs.
Try it · concept lab
How many GPUs does a model need?
Pick a model size and a precision.
Precision: bytes per number
Memory per GPU
Weights
140 GB
GPUs to answer
3
Training memory
1.1 TB
GPUs just to fit training
16
Rules of thumb: serving ≈ weights + ~20% working memory; training ≈ 16 bytes per parameter (weights, gradients, optimizer). Real labs use far more GPUs than this, for speed.
Big idea: Model size × bytes per number = memory. Shrinking the numbers (lower precision) is one of the biggest tricks in AI hardware.
Swap it: other ways to do the same job
Data parallel
Every GPU holds the whole model and gets different data. Simple, but the model must fit on one GPU.
Tensor / pipeline parallel
Slice the model itself across GPUs, by layer or within a layer. Fits giant models, needs very fast links.
Fine-tuning
Start from an existing model and train a little more on your data. Hours on a few GPUs instead of months on thousands.
Go deeper
What does learning cost?
Training the biggest models can take thousands of chips running for weeks or months, and the electricity alone can cost millions of dollars. Answering everyone's questions afterward can use even more. Some people say it is worth it, because one model can help millions of people. Others say that energy, and the pollution from making it, could do more good elsewhere. Smaller models that use less energy keep improving.
Figures are rough, as of 2026.
Talk about it
- Q1
Thousands of GPUs work for weeks to teach one model. What have you learned that took weeks of practice?
- Q2
Starting from a model someone else trained takes hours instead of months. When is it fair to build on someone else's work?
For grown-ups: there are no right answers here. Ask a question, then ask "why do you think that?" The reasons matter more than the answer.
Printable question sheet (PDF)