Skip to content
Hi, Bot

The AI Server

Eight GPUs, two CPUs, eight network cards, one very loud box.

How big?Like a large suitcase full of metal. Often over 100 kg.

Real: What it looks like.

Keys: arrows rotate · + / − zoom · 0 reset · 1–4 views · S signal · T tour · L labels · Space spin. Models are stylised and built from code: proportions are honest, details are simplified.

What it is

The top tray holds eight accelerator modules on a baseboard, linked by switch chips so any GPU can reach any other at very high speed.

The bottom tray is a more normal computer: two CPUs, terabytes of memory, fast SSDs. Eight network cards, one per GPU, connect to the rest of the cluster.

A wall of fans pulls air from front to back, or, in liquid-cooled versions, hoses carry the heat away.

Why AI needs it

This box is the basic building block of AI clusters. Training runs are counted in thousands of them; each one costs about as much as several houses.

Every labeled part

  1. 1

    8 accelerators

    Under tall heat sinks. Signal mode shows them working.

  2. 2

    GPU switch chips

    Let all eight GPUs talk to each other at once, much faster than the network.

  3. 3

    Network cards

    One per GPU, out the back to the cluster.

  4. 4

    2 CPUs

    Coordinate the work and feed the GPUs.

  5. 5

    Memory sticks

    Terabytes of system DRAM for staging data.

  6. 6

    Power supplies

    Several, so one can fail safely.

  7. 7

    Fan wall

    Air in the front, heat out the back. Watch the blue airflow in Signal mode.

  8. 8

    NVMe drives

    Local fast storage for data and checkpoints.

  9. 9

    Lid

    Explode to lift it; trays separate so you can see each layer.

Try it · concept lab

How many GPUs does a model need?

Pick a model size and a precision.

Precision: bytes per number

Memory per GPU

Weights

140 GB

GPUs to answer

3

Training memory

1.1 TB

GPUs just to fit training

16

Rules of thumb: serving ≈ weights + ~20% working memory; training ≈ 16 bytes per parameter (weights, gradients, optimizer). Real labs use far more GPUs than this, for speed.

Big idea: Model size × bytes per number = memory. Shrinking the numbers (lower precision) is one of the biggest tricks in AI hardware.

Try it · concept lab

Fill a rack

Add servers until something gives.

Rack type

dashed line = what the rack can power and cool

Rack draw

43 kW

≈ US homes (avg)

36

Verdict

fits

Assumes ~10 kW per 8-GPU server plus ~2.5 kW for switches and fans; an average US home uses about 1.2 kW around the clock. Rough figures.

Big idea: AI racks are limited by power and cooling, not space. That's why liquid cooling and new power designs matter.

Swap it: other ways to do the same job

  • 4-GPU or 1-GPU server

    Smaller and cheaper. Common for running models (inference) rather than training them.

  • Rack-scale system

    Dozens of GPUs across a whole rack wired so tightly they act like one giant GPU.

  • Workstation

    One or two GPUs under a desk. Great for learning and small models.

Talk about it

  1. Q1

    This box holds eight GPUs and costs about as much as several houses. Why do you think it costs so much?

  2. Q2

    A small computer with one GPU is great for learning. What would you build if you had one at home?

For grown-ups: there are no right answers here. Ask a question, then ask "why do you think that?" The reasons matter more than the answer.

Printable question sheet (PDF)