The Software Stack
Eight layers between 'hi' and a transistor flipping.
How big?Millions of lines of code, layered like a cake.
Real: What it looks like.
Keys: arrows rotate · + / − zoom · 0 reset · 1–4 views · S signal · T tour · L labels · Space spin. Models are stylised and built from code: proportions are honest, details are simplified.
What it is
At the top is the app you see. Under it, serving software queues requests and spreads them over GPUs. The model is a file of weights plus a recipe for using them.
ML frameworks (like PyTorch or JAX) describe the math. Compilers and kernels (like CUDA, Triton or XLA) turn that math into instructions tuned for a specific chip. Drivers and the runtime move memory and launch work. The operating system and container schedulers keep thousands of machines organized.
At the bottom is everything else in this atlas.
Why AI needs it
The same chip can be twice as fast with better software. A lot of AI progress is people finding cleverer ways to keep the hardware busy.
Every labeled part
- 1
App
The chat window. It only sees text in and text out.
- 2
Serving
Batches requests, picks GPUs, streams tokens back.
- 3
Model
Architecture plus billions of weights, often hundreds of gigabytes.
- 4
Framework
Where researchers write models as math on tensors (grids of numbers).
- 5
Kernels + compilers
Hand-tuned or compiled routines for each operation on each chip.
- 6
Drivers + runtime
Copy memory to the GPU, launch kernels, report errors.
- 7
OS + containers
Linux plus schedulers that place jobs on thousands of machines.
- 8
Hardware
Chips, memory, network, power, cooling.
Try it · concept lab
Walk down the stack
Tap each layer to see what it sees.
What the app layer sees
"Why is the sky blue?"
Token numbers and sizes are made up for illustration; the shapes are realistic.
Big idea: Each layer hides the one below it. The app sees words; the hardware sees voltages.
Swap it: other ways to do the same job
PyTorch vs. JAX
PyTorch is flexible and dominant in research; JAX compiles whole programs and shines on TPUs.
Hand-written kernels vs. compilers
Hand-tuned is fastest for one chip; compilers are portable and quicker to write.
Managed API vs. self-hosting
Call someone else's model over the internet, or run open weights on your own GPUs.
Talk about it
- Q1
Better software can make the same chip about twice as fast. When has a better method made you faster at something?
- Q2
You can use someone else's AI over the internet, or run your own. What are good reasons to pick each one?
For grown-ups: there are no right answers here. Ask a question, then ask "why do you think that?" The reasons matter more than the answer.
Printable question sheet (PDF)