The AI Server
Eight GPUs, two CPUs, eight network cards, one very loud box.
How big?Like a large suitcase full of metal. Often over 100 kg.
Real: What it looks like.
Keys: arrows rotate · + / − zoom · 0 reset · 1–4 views · S signal · T tour · L labels · Space spin. Models are stylised and built from code: proportions are honest, details are simplified.
What it is
The top tray holds eight accelerator modules on a baseboard, linked by switch chips so any GPU can reach any other at very high speed.
The bottom tray is a more normal computer: two CPUs, terabytes of memory, fast SSDs. Eight network cards, one per GPU, connect to the rest of the cluster.
A wall of fans pulls air from front to back, or, in liquid-cooled versions, hoses carry the heat away.
Why AI needs it
This box is the basic building block of AI clusters. Training runs are counted in thousands of them; each one costs about as much as several houses.
Every labeled part
- 1
8 accelerators
Under tall heat sinks. Signal mode shows them working.
- 2
GPU switch chips
Let all eight GPUs talk to each other at once, much faster than the network.
- 3
Network cards
One per GPU, out the back to the cluster.
- 4
2 CPUs
Coordinate the work and feed the GPUs.
- 5
Memory sticks
Terabytes of system DRAM for staging data.
- 6
Power supplies
Several, so one can fail safely.
- 7
Fan wall
Air in the front, heat out the back. Watch the blue airflow in Signal mode.
- 8
NVMe drives
Local fast storage for data and checkpoints.
- 9
Lid
Explode to lift it; trays separate so you can see each layer.
Try it · concept lab
How many GPUs does a model need?
Pick a model size and a precision.
Precision: bytes per number
Memory per GPU
Weights
140 GB
GPUs to answer
3
Training memory
1.1 TB
GPUs just to fit training
16
Rules of thumb: serving ≈ weights + ~20% working memory; training ≈ 16 bytes per parameter (weights, gradients, optimizer). Real labs use far more GPUs than this, for speed.
Big idea: Model size × bytes per number = memory. Shrinking the numbers (lower precision) is one of the biggest tricks in AI hardware.
Try it · concept lab
Fill a rack
Add servers until something gives.
Rack type
dashed line = what the rack can power and cool
Rack draw
43 kW
≈ US homes (avg)
36
Verdict
fits
Assumes ~10 kW per 8-GPU server plus ~2.5 kW for switches and fans; an average US home uses about 1.2 kW around the clock. Rough figures.
Big idea: AI racks are limited by power and cooling, not space. That's why liquid cooling and new power designs matter.
Swap it: other ways to do the same job
4-GPU or 1-GPU server
Smaller and cheaper. Common for running models (inference) rather than training them.
Rack-scale system
Dozens of GPUs across a whole rack wired so tightly they act like one giant GPU.
Workstation
One or two GPUs under a desk. Great for learning and small models.
Talk about it
- Q1
This box holds eight GPUs and costs about as much as several houses. Why do you think it costs so much?
- Q2
A small computer with one GPU is great for learning. What would you build if you had one at home?
For grown-ups: there are no right answers here. Ask a question, then ask "why do you think that?" The reasons matter more than the answer.
Printable question sheet (PDF)