Skip to content
Hi, Bot

The Software Stack

Eight layers between 'hi' and a transistor flipping.

How big?Millions of lines of code, layered like a cake.

Real: What it looks like.

Keys: arrows rotate · + / − zoom · 0 reset · 1–4 views · S signal · T tour · L labels · Space spin. Models are stylised and built from code: proportions are honest, details are simplified.

What it is

At the top is the app you see. Under it, serving software queues requests and spreads them over GPUs. The model is a file of weights plus a recipe for using them.

ML frameworks (like PyTorch or JAX) describe the math. Compilers and kernels (like CUDA, Triton or XLA) turn that math into instructions tuned for a specific chip. Drivers and the runtime move memory and launch work. The operating system and container schedulers keep thousands of machines organized.

At the bottom is everything else in this atlas.

Why AI needs it

The same chip can be twice as fast with better software. A lot of AI progress is people finding cleverer ways to keep the hardware busy.

Every labeled part

  1. 1

    App

    The chat window. It only sees text in and text out.

  2. 2

    Serving

    Batches requests, picks GPUs, streams tokens back.

  3. 3

    Model

    Architecture plus billions of weights, often hundreds of gigabytes.

  4. 4

    Framework

    Where researchers write models as math on tensors (grids of numbers).

  5. 5

    Kernels + compilers

    Hand-tuned or compiled routines for each operation on each chip.

  6. 6

    Drivers + runtime

    Copy memory to the GPU, launch kernels, report errors.

  7. 7

    OS + containers

    Linux plus schedulers that place jobs on thousands of machines.

  8. 8

    Hardware

    Chips, memory, network, power, cooling.

Try it · concept lab

Walk down the stack

Tap each layer to see what it sees.

What the app layer sees

"Why is the sky blue?"

Token numbers and sizes are made up for illustration; the shapes are realistic.

Big idea: Each layer hides the one below it. The app sees words; the hardware sees voltages.

Swap it: other ways to do the same job

  • PyTorch vs. JAX

    PyTorch is flexible and dominant in research; JAX compiles whole programs and shines on TPUs.

  • Hand-written kernels vs. compilers

    Hand-tuned is fastest for one chip; compilers are portable and quicker to write.

  • Managed API vs. self-hosting

    Call someone else's model over the internet, or run open weights on your own GPUs.

Talk about it

  1. Q1

    Better software can make the same chip about twice as fast. When has a better method made you faster at something?

  2. Q2

    You can use someone else's AI over the internet, or run your own. What are good reasons to pick each one?

For grown-ups: there are no right answers here. Ask a question, then ask "why do you think that?" The reasons matter more than the answer.

Printable question sheet (PDF)