Skip to content
Hi, Bot

Hi, Bot · First Principles · Phase 6: Build it

Lesson 48 of 48

Reinforcement learning

Learning from consequences instead of answers.

No label says what you should have done — only a reward saying how it went, often much later. That delay is the entire difficulty, and the reason credit assignment is the field's central problem.

Do this

Build a gridworld with a goal, a pit and walls. Solve it twice: once with tabular Q-learning, once with a policy gradient using the network you wrote at lesson 45.

The question that unlocks the next lesson

How does reinforcement learning differ from supervised learning?

  • AIt uses more data
  • BThere is no correct-answer label — only a reward signal, often delayed, that must be assigned back to earlier actions
  • CIt never uses gradients
  • DIt only works on games

Start at lesson 1 and work up to this one

48 lessons, one a day. Answer each lesson's question correctly and the next one opens immediately — nothing here is unlocked by waiting.

By submitting, you agree to our Terms and Privacy Policy.