Lesson 41 of 48
Convexity and gradient descent
Rolling downhill, and when you can trust where you land.
On a convex surface, any local minimum is the global one and descent is trustworthy. Deep networks are not convex, which is why training is an empirical craft with a mathematical core.
Do this
Implement gradient descent from scratch on a convex function and show it reaches the minimum. Then run it on a non-convex one from several starting points and compare where it lands.
The question that unlocks the next lesson
Why does convexity matter for gradient descent?
- AIt makes each step faster to compute
- BOn a convex function every local minimum is the global minimum, so descent cannot get stuck in a worse one
- CIt guarantees the gradient is always zero
- DIt removes the need to choose a learning rate