Your Prompt's Round Trip
What happens in the second after you hit send.
How big?Hundreds to thousands of kilometres, a fraction of a second each way.
Real: What it looks like.
Keys: arrows rotate · + / − zoom · 0 reset · 1–4 views · S signal · T tour · L labels · Space spin. Models are stylised and built from code: proportions are honest, details are simplified.
What it is
Your question leaves your phone as radio, hops onto fiber at a tower or router, and crosses the internet to a data center.
A load balancer picks a GPU server that has the model loaded. The GPU reads your whole prompt at once (prefill), then generates the answer one token at a time (decode), each new token depending on all the ones before it.
A KV cache keeps the math for earlier words so they aren't recomputed. Tokens stream back to you as they're made, which is why answers appear word by word.
Why AI needs it
Training happens once. Inference happens every time anyone asks anything, billions of times a day. Most AI hardware in the world is now serving answers.
Every labeled part
- 1
Your device
Your words become packets.
- 2
Wi-Fi / cell tower
Radio to fiber.
- 3
The internet
Fiber across cities and under oceans. Light covers about 200 km every millisecond.
- 4
Load balancer
Sends you to a server with spare room. Many requests are batched together on one GPU.
- 5
GPU server
Prefill reads your prompt in parallel; decode writes the answer token by token.
- 6
KV cache
Memory of the conversation so far. Long chats need lots of HBM.
- 7
Tokens back
Light-blue pulses: the answer streaming home, piece by piece.
Try it · concept lab
Where does the wait come from?
Move the sliders and watch the timeline.
- Network 45 ms
- Queue 50 ms
- Read prompt (prefill) 50 ms
- Write answer (decode) 4.2 s
First word appears
145 ms
Whole answer
4.3 s
Illustrative rates: prefill ≈ 8,000 tokens/s, decode ≈ 60 tokens/s, light in fiber ≈ 200 km/ms. Writing the answer usually takes longest, because each word waits for the one before it.
Big idea: Distance costs milliseconds, but writing the answer one token at a time is usually the longest part.
Swap it: other ways to do the same job
On-device
A small model on your phone's NPU. Private and offline, but less capable.
Edge
Servers in your city instead of a far-off campus. Faster round trip, smaller models.
Speculative decoding
A small model drafts several tokens and the big model checks them in one pass. Same answer, faster.
Talk about it
- Q1
Your question travels to a far-off building and back in about a second. What else do you know that moves that fast?
- Q2
A small AI can run on your phone without sending anything away. When would you want that instead of a big one?
For grown-ups: there are no right answers here. Ask a question, then ask "why do you think that?" The reasons matter more than the answer.
Printable question sheet (PDF)