The Accelerator Module
GPU, memory, power delivery and a cold plate on one board.
How big?About the size of a large paperback book. It can draw as much power as a space heater.
Real: What it looks like.
Keys: arrows rotate · + / − zoom · 0 reset · 1–4 views · S signal · T tour · L labels · Space spin. Models are stylised and built from code: proportions are honest, details are simplified.
What it is
The GPU die and its HBM stacks sit on an interposer, which sits on a package substrate, which sits on a module board. Each layer fans the wiring out to something coarser.
Dozens of voltage regulators (the grey blocks) turn the server's power into the very low voltage and very high current the chip needs.
A copper cold plate presses on top. Liquid flowing through tiny channels inside carries the heat away.
Why AI needs it
This module is the unit you count when people say a lab trained on '10,000 GPUs'. It is also where most of the electricity and heat in an AI data center ends up.
Every labeled part
- 1
GPU die
The compute. Open the GPU Die exhibit to go inside it.
- 2
HBM stacks
Six stacks of memory hugging the die.
- 3
Interposer
Silicon bridge carrying thousands of fine wires between die and memory.
- 4
Package substrate
Spreads those wires out to a scale the circuit board can handle.
- 5
Voltage regulators
They step power down to around 1 volt at hundreds of amps. Watch the yellow power flow in Signal mode.
- 6
Cold plate
Solid copper with liquid running through it. Explode the view to lift it off; X-ray it to see the micro-fins inside.
- 7
Coolant in / out
Cool liquid enters (blue), picks up heat, and leaves warm (red).
- 8
Mezzanine connectors
Underneath, high-density connectors plug the module into the server's baseboard.
Try it · concept lab
How many GPUs does a model need?
Pick a model size and a precision.
Precision: bytes per number
Memory per GPU
Weights
140 GB
GPUs to answer
3
Training memory
1.1 TB
GPUs just to fit training
16
Rules of thumb: serving ≈ weights + ~20% working memory; training ≈ 16 bytes per parameter (weights, gradients, optimizer). Real labs use far more GPUs than this, for speed.
Big idea: Model size × bytes per number = memory. Shrinking the numbers (lower precision) is one of the biggest tricks in AI hardware.
Swap it: other ways to do the same job
PCIe card
Plugs into a standard slot, like a gaming card. Easy to deploy, but lower power and slower links between GPUs.
Mezzanine module (SXM / OAM style)
Lies flat on a baseboard. More power, faster GPU-to-GPU links, needs a special server.
CPU + GPU 'superchip'
Puts a CPU and GPU on one board with a very fast link between them, so the CPU's memory acts like extra GPU memory.
Talk about it
- Q1
People say a lab trained on 10,000 GPUs. How much space do you think that many boards would take?
- Q2
One of these modules uses about as much power as a microwave running all the time. Where should that electricity come from?
For grown-ups: there are no right answers here. Ask a question, then ask "why do you think that?" The reasons matter more than the answer.
Printable question sheet (PDF)