Skip to content
Hi, Bot

HBM: Stacked Memory

Memory chips stacked like pancakes and drilled through with copper.

How big?Each stack is about the size of a fingernail clipping and a little taller than a coin is thick.

Real: What it looks like.

Keys: arrows rotate · + / − zoom · 0 reset · 1–4 views · S signal · T tour · L labels · Space spin. Models are stylised and built from code: proportions are honest, details are simplified.

What it is

High-bandwidth memory is several DRAM dies stacked on top of a logic die. Thousands of copper columns called through-silicon vias (TSVs) run straight down through the stack.

The stack sits right beside the GPU on a silicon interposer, a slab of silicon with very fine wiring. The connection is extremely wide: a thousand-plus wires per stack instead of a few dozen.

Wide plus short means fast. That is the whole trick.

Why AI needs it

An AI model's weights have to stream from memory into the cores constantly. Often the GPU is waiting on memory rather than math, so memory bandwidth decides how fast a model runs.

Every labeled part

  1. 1

    The stack

    8 to 16 memory dies plus a base die, sealed in black molding compound. Explode to fan the layers apart.

  2. 2

    DRAM die

    Each layer is a full memory chip, thinned to a fraction of a hair's width so a tall stack still fits.

  3. 3

    Base logic die

    Directs traffic between the stacked dies and the GPU.

  4. 4

    Through-silicon vias

    Copper columns drilled through every die, so data goes straight up and down the stack. Visible in X-ray.

  5. 5

    Micro-bumps

    Microscopic solder joints connecting the stack to the interposer, many thousands of them.

  6. 6

    Silicon interposer

    A thin silicon slab that carries extremely fine wires between memory and GPU, far finer than a circuit board can.

  7. 7

    Interposer wiring

    The short, wide data highway. Turn on Signal to watch data race from the stack to the GPU.

  8. 8

    GPU edge

    The GPU die sits millimetres away so the trip is as short as possible.

Try it · concept lab

Pipes, not just processors

Drag the model size and compare how long each pipe takes.

  • HBM (whole GPU)35 ms
  • GPU ↔ GPU link156 ms
  • Server DRAM467 ms
  • 800 Gb/s network1.4 s
  • PCIe 5 slot2.2 s
  • NVMe SSD20 s
  • Home internet (100 Mb/s)3.1 h

Max tokens/s, one user (HBM)

≈ 29

HBM vs home internet

320,000×

To write each new word, a GPU reads (roughly) every weight once. So for one user, speed ≈ memory bandwidth ÷ model size. Rough 2026 figures.

Big idea: For chat, a GPU usually waits on memory, not math. That's why HBM is stacked right next to the die.

Swap it: other ways to do the same job

  • GDDR memory

    Separate chips on the circuit board, used in gaming GPUs. Cheaper and easier to make, with less bandwidth.

  • Regular DRAM (DDR)

    Huge capacity and cheap per gigabyte, but much slower. This is the server's main memory.

  • On-chip SRAM

    Fastest of all and right inside the compute die, but tiny and expensive per byte.

Talk about it

  1. Q1

    Stacked memory works because the data's trip is short. Where in your life does a shorter trip save lots of time?

  2. Q2

    Often the chip sits waiting for memory, not doing math. Is a fast worker useful if their supplies arrive slowly?

For grown-ups: there are no right answers here. Ask a question, then ask "why do you think that?" The reasons matter more than the answer.

Printable question sheet (PDF)