HBM: Stacked Memory
Memory chips stacked like pancakes and drilled through with copper.
How big?Each stack is about the size of a fingernail clipping and a little taller than a coin is thick.
Real: What it looks like.
Keys: arrows rotate · + / − zoom · 0 reset · 1–4 views · S signal · T tour · L labels · Space spin. Models are stylised and built from code: proportions are honest, details are simplified.
What it is
High-bandwidth memory is several DRAM dies stacked on top of a logic die. Thousands of copper columns called through-silicon vias (TSVs) run straight down through the stack.
The stack sits right beside the GPU on a silicon interposer, a slab of silicon with very fine wiring. The connection is extremely wide: a thousand-plus wires per stack instead of a few dozen.
Wide plus short means fast. That is the whole trick.
Why AI needs it
An AI model's weights have to stream from memory into the cores constantly. Often the GPU is waiting on memory rather than math, so memory bandwidth decides how fast a model runs.
Every labeled part
- 1
The stack
8 to 16 memory dies plus a base die, sealed in black molding compound. Explode to fan the layers apart.
- 2
DRAM die
Each layer is a full memory chip, thinned to a fraction of a hair's width so a tall stack still fits.
- 3
Base logic die
Directs traffic between the stacked dies and the GPU.
- 4
Through-silicon vias
Copper columns drilled through every die, so data goes straight up and down the stack. Visible in X-ray.
- 5
Micro-bumps
Microscopic solder joints connecting the stack to the interposer, many thousands of them.
- 6
Silicon interposer
A thin silicon slab that carries extremely fine wires between memory and GPU, far finer than a circuit board can.
- 7
Interposer wiring
The short, wide data highway. Turn on Signal to watch data race from the stack to the GPU.
- 8
GPU edge
The GPU die sits millimetres away so the trip is as short as possible.
Try it · concept lab
Pipes, not just processors
Drag the model size and compare how long each pipe takes.
- HBM (whole GPU)35 ms
- GPU ↔ GPU link156 ms
- Server DRAM467 ms
- 800 Gb/s network1.4 s
- PCIe 5 slot2.2 s
- NVMe SSD20 s
- Home internet (100 Mb/s)3.1 h
Max tokens/s, one user (HBM)
≈ 29
HBM vs home internet
320,000×
To write each new word, a GPU reads (roughly) every weight once. So for one user, speed ≈ memory bandwidth ÷ model size. Rough 2026 figures.
Big idea: For chat, a GPU usually waits on memory, not math. That's why HBM is stacked right next to the die.
Swap it: other ways to do the same job
GDDR memory
Separate chips on the circuit board, used in gaming GPUs. Cheaper and easier to make, with less bandwidth.
Regular DRAM (DDR)
Huge capacity and cheap per gigabyte, but much slower. This is the server's main memory.
On-chip SRAM
Fastest of all and right inside the compute die, but tiny and expensive per byte.
Talk about it
- Q1
Stacked memory works because the data's trip is short. Where in your life does a shorter trip save lots of time?
- Q2
Often the chip sits waiting for memory, not doing math. Is a fast worker useful if their supplies arrive slowly?
For grown-ups: there are no right answers here. Ask a question, then ask "why do you think that?" The reasons matter more than the answer.
Printable question sheet (PDF)