Lesson 48 of 48
Reinforcement learning
Learning from consequences instead of answers.
No label says what you should have done — only a reward saying how it went, often much later. That delay is the entire difficulty, and the reason credit assignment is the field's central problem.
Do this
Build a gridworld with a goal, a pit and walls. Solve it twice: once with tabular Q-learning, once with a policy gradient using the network you wrote at lesson 45.
The question that unlocks the next lesson
How does reinforcement learning differ from supervised learning?
- AIt uses more data
- BThere is no correct-answer label — only a reward signal, often delayed, that must be assigned back to earlier actions
- CIt never uses gradients
- DIt only works on games