Skip to content
Hi, Bot

A Training Run

Thousands of GPUs learning one model together for weeks.

How big?Weeks to months, thousands to 100,000+ GPUs.

Real: What it looks like.

Keys: arrows rotate · + / − zoom · 0 reset · 1–4 views · S signal · T tour · L labels · Space spin. Models are stylised and built from code: proportions are honest, details are simplified.

What it is

Data streams from the lake through CPU data loaders into the GPUs in batches. Each GPU guesses, measures how wrong it was, and works out how to nudge the model's weights to be less wrong.

Then comes the all-reduce: every GPU shares its nudges over the network and they're averaged, so all copies of the model stay identical. This happens for every step, millions of times.

Every few hours the run saves a checkpoint. If a GPU dies (and at this scale one always does), training restarts from the last save.

Why AI needs it

This is how a model like a chatbot gets made. The hardware's job is to keep every GPU busy and never waiting on data, memory or the network.

Every labeled part

  1. 1

    Data lake

    Gold-zone training data.

  2. 2

    Data loaders

    CPUs read, shuffle and batch the data so GPUs never starve.

  3. 3

    GPU servers

    Each runs a copy (or a slice) of the model on its batch.

  4. 4

    Network fabric

    Leaf and spine switches. Explode to lift them above the nodes.

  5. 5

    All-reduce ring

    Light-blue pulses: every GPU passing its share of the update around the ring until all agree.

  6. 6

    The model

    Billions of weights, slowly getting better at predicting.

  7. 7

    Checkpoints

    Saved progress, so a failure costs hours, not weeks.

Try it · concept lab

Agreeing on an answer: the ring

Pass numbers around until every GPU knows the total.

GPU 13GPU 27GPU 31GPU 49GPU 54GPU 66everyone needssum = 30

Passes

0 / 10

Each GPU only ever talks to its neighbour, so no single link gets swamped. Real all-reduce splits the numbers into chunks and does both phases on every link at once; the idea is the same.

Big idea: After every training step, all GPUs must agree. Passing around a ring keeps every link equally busy, so it scales to thousands of GPUs.

Try it · concept lab

How many GPUs does a model need?

Pick a model size and a precision.

Precision: bytes per number

Memory per GPU

Weights

140 GB

GPUs to answer

3

Training memory

1.1 TB

GPUs just to fit training

16

Rules of thumb: serving ≈ weights + ~20% working memory; training ≈ 16 bytes per parameter (weights, gradients, optimizer). Real labs use far more GPUs than this, for speed.

Big idea: Model size × bytes per number = memory. Shrinking the numbers (lower precision) is one of the biggest tricks in AI hardware.

Swap it: other ways to do the same job

  • Data parallel

    Every GPU holds the whole model and gets different data. Simple, but the model must fit on one GPU.

  • Tensor / pipeline parallel

    Slice the model itself across GPUs, by layer or within a layer. Fits giant models, needs very fast links.

  • Fine-tuning

    Start from an existing model and train a little more on your data. Hours on a few GPUs instead of months on thousands.

Go deeper

What does learning cost?

Training the biggest models can take thousands of chips running for weeks or months, and the electricity alone can cost millions of dollars. Answering everyone's questions afterward can use even more. Some people say it is worth it, because one model can help millions of people. Others say that energy, and the pollution from making it, could do more good elsewhere. Smaller models that use less energy keep improving.

Figures are rough, as of 2026.

Talk about it

  1. Q1

    Thousands of GPUs work for weeks to teach one model. What have you learned that took weeks of practice?

  2. Q2

    Starting from a model someone else trained takes hours instead of months. When is it fair to build on someone else's work?

For grown-ups: there are no right answers here. Ask a question, then ask "why do you think that?" The reasons matter more than the answer.

Printable question sheet (PDF)