Lesson 46 of 48
Attention from scratch
A dot product decides what to look at, and everything follows from that.
The mechanism behind every model you have used is a similarity score, a softmax, and a weighted average. Implementing it once removes the mystique permanently.
Do this
Implement scaled dot-product attention in NumPy. On a short sequence you design yourself, show that each output row is a weighted average of value rows with weights you can predict.
The question that unlocks the next lesson
In scaled dot-product attention, what do the query–key dot products become after the softmax?
- AThe output vectors themselves
- BWeights that sum to one, used to average the value vectors
- CThe gradients for the next layer
- DThe positional encodings