Skip to content
Hi, Bot

Hi, Bot · First Principles · Phase 5: Spectra, probability, optimization

Lesson 41 of 48

Convexity and gradient descent

Rolling downhill, and when you can trust where you land.

On a convex surface, any local minimum is the global one and descent is trustworthy. Deep networks are not convex, which is why training is an empirical craft with a mathematical core.

Do this

Implement gradient descent from scratch on a convex function and show it reaches the minimum. Then run it on a non-convex one from several starting points and compare where it lands.

The question that unlocks the next lesson

Why does convexity matter for gradient descent?

  • AIt makes each step faster to compute
  • BOn a convex function every local minimum is the global minimum, so descent cannot get stuck in a worse one
  • CIt guarantees the gradient is always zero
  • DIt removes the need to choose a learning rate

Start at lesson 1 and work up to this one

48 lessons, one a day. Answer each lesson's question correctly and the next one opens immediately — nothing here is unlocked by waiting.

By submitting, you agree to our Terms and Privacy Policy.