The GPU Die
Thousands of small cores doing the same math at the same time.
How big?About the size of a postage stamp, and roughly as big as a chip can be printed in one shot.
Real: What it looks like.
Keys: arrows rotate · + / − zoom · 0 reset · 1–4 views · S signal · T tour · L labels · Space spin. Models are stylised and built from code: proportions are honest, details are simplified.
What it is
A GPU (graphics processing unit) was built to color millions of pixels at once. It turns out neural networks need the same thing: the same simple math done to huge grids of numbers in parallel.
The die is divided into many identical blocks called streaming multiprocessors (SMs). Each has dozens of simple cores plus tensor cores, which are circuits that do a whole small matrix multiply in one step.
A big strip of on-chip cache (L2) sits in the middle. Memory controllers line the edges, wired to the HBM stacks next door.
Why AI needs it
Training and running AI is mostly matrix multiplication. A CPU does a few big jobs quickly; a GPU does thousands of small jobs at once. For AI, the GPU's way wins.
Every labeled part
- 1
Streaming multiprocessor
One of ~100+ identical compute blocks. Each runs many threads at once. Turn on Signal to watch them light up as they work.
- 2
Tensor core
A tiny matrix-multiply machine. It takes small grids of numbers and multiplies-and-adds them in one go. This is where most AI math happens.
- 3
L2 cache
Fast on-chip memory shared by all the SMs. Keeping data here avoids slow trips to off-chip memory.
- 4
Memory controllers
The edge of the die that talks to the stacked HBM memory. Thousands of wires leave from here.
- 5
Chip-to-chip links
High-speed connections to other GPUs so eight (or many more) can share work as if they were one big chip.
- 6
Metal layers
What a photo of a die actually shows: the top of a dense stack of copper wiring. X-ray makes it glass.
- 7
Solder bumps
Thousands of tiny solder balls under the die carry power and signals down into the package.
Try it · concept lab
CPU vs. GPU: the race
Same 256 multiplications, two very different chips.
CPU · 8 fast cores
0/256 multiplications · 0 ticks
GPU · 256 slower cores
0/256 · done in 4 ticks
Each GPU core is 4× slower here, and it still wins 8× because it does them all at once. Real chips differ in the details; the shape of the result is right.
Big idea: AI math is lots of small, independent pieces. Many slow workers beat a few fast ones when the jobs don't depend on each other.
Swap it: other ways to do the same job
CPU
A few powerful, flexible cores. Great at branching logic and the operating system; far slower at giant grids of math.
TPU / custom AI chip (ASIC)
Built only for matrix math, so it is more efficient at it, but less flexible if AI methods change.
Wafer-scale chip
Keeps the whole silicon wafer as one enormous chip. Huge on-chip memory and speed, but exotic to cool, power and build.
Chiplets
Several smaller dies linked side by side. Gets past the size limit and wastes less silicon, but adds link overhead.
Go deeper
Who makes the chips?
The most advanced AI chips are built in only a handful of factories, mostly in East Asia. Each factory costs tens of billions of dollars. Having so few lets them get good at it, but a storm, an earthquake or a trade fight in one place can slow down AI everywhere. Some countries are paying to build factories at home, and people disagree about whether that is worth the cost.
Figures are rough, as of 2026.
Talk about it
- Q1
A GPU does thousands of small jobs at once. What chore at home goes faster when lots of people help at once?
- Q2
Is there a job where lots of helpers would actually make things slower? Why?
- Q3
Most big AI chips come from just a few factories. Is that risky, or is it a smart way to do it?
For grown-ups: there are no right answers here. Ask a question, then ask "why do you think that?" The reasons matter more than the answer.
Printable question sheet (PDF)