Skip to content
Hi, Bot

Network Switch

The traffic cop: dozens of ports, one giant routing chip.

How big?A flat box as wide as a rack, about as tall as two stacked phones.

Real: What it looks like.

Keys: arrows rotate · + / − zoom · 0 reset · 1–4 views · S signal · T tour · L labels · Space spin. Models are stylised and built from code: proportions are honest, details are simplified.

What it is

A switch reads where each packet is going and sends it out of the right port, billions of times a second. One big chip (the switch ASIC) does it, under the biggest heat sink in the box.

AI clusters wire switches in layers: 'leaf' switches connect to servers, 'spine' switches connect leaves together. Any GPU can then reach any other in a few hops.

Why AI needs it

When thousands of GPUs average their results after every training step, every one of them talks at once. The network has to carry all of it without becoming the bottleneck.

Every labeled part

  1. 1

    Ports

    32–64 high-speed ports on the front. Signal mode routes packets between them.

  2. 2

    Switch chip

    Decides where every packet goes. Can move tens of terabits per second.

  3. 3

    Fans

    Pull air front to back across the chip.

  4. 4

    Power supplies

    Two of them, so one can fail without the switch going down.

  5. 5

    Lid

    Explode to lift it off; X-ray to see straight through.

Try it · concept lab

Agreeing on an answer: the ring

Pass numbers around until every GPU knows the total.

GPU 13GPU 27GPU 31GPU 49GPU 54GPU 66everyone needssum = 30

Passes

0 / 10

Each GPU only ever talks to its neighbour, so no single link gets swamped. Real all-reduce splits the numbers into chunks and does both phases on every link at once; the idea is the same.

Big idea: After every training step, all GPUs must agree. Passing around a ring keeps every link equally busy, so it scales to thousands of GPUs.

Swap it: other ways to do the same job

  • Leaf-spine

    The standard layout. Every server is a few hops from every other.

  • Rail-optimized

    GPU #1 in every server shares a switch, GPU #2 another, and so on. Fewer hops for AI traffic patterns.

  • Direct links / torus

    Chips wire straight to neighbours (as in TPU pods), sometimes through optical circuit switches that physically redirect light with mirrors.

Talk about it

  1. Q1

    The switch is a traffic cop for data. What makes a good traffic cop on a busy road?

  2. Q2

    If every chip talks at the same time, the network jams. What rules would you make so everyone gets a turn?

For grown-ups: there are no right answers here. Ask a question, then ask "why do you think that?" The reasons matter more than the answer.

Printable question sheet (PDF)