Gradient descent: rolling downhill
How models find good weights: feel the slope, take a small step down, repeat.
Picture yourself on a foggy hillside at night. You can't see the valley, but you can feel which way the ground slopes under your feet. Step downhill. Feel again. Step again. That's gradient descent, the algorithm that trains almost every model in existence.
๐ The gradient
The gradient is the slope of the loss with respect to each weight. It points uphill.
So we move in the opposite direction: w_new = w_old โ learning_rate ร gradient.
๐ The learning rate
The learning rate is the step size.
Too small: you'll get there, eventually, in a few weeks. Too big: you leap over the valley and bounce around or fly off entirely.
Picking it well is one of the most important practical skills in deep learning.
Training loss is jumping up and down wildly and getting worse. What's the most likely fix?
๐ Mini-batches (SGD)
Computing the gradient on the whole dataset each step is slow. Instead we use a small random mini-batch (say 32 or 256 examples).
This is stochastic gradient descent (SGD). The noisy steps are faster, and the noise even helps escape bad spots.
๐ Fancier optimisers
Momentum keeps rolling in a consistent direction, like a ball gathering speed.
Adam adapts the step size for every weight individually. It's the default choice for most deep learning today.
Which way does gradient descent step?
โจ Before you drift off
- Gradient = slope of the loss; it points uphill.
- Update: w โ w โ learning_rate ร gradient.
- Learning rate too high โ bouncing; too low โ crawling.
- Mini-batches (SGD) and Adam make it fast in practice.
๐ Go deeper (free & open)
- Gradient descent, how neural networks learn โ ยท 3Blue1Brown
- Why Momentum Really Works โ ยท Gabriel Goh, Distill ยท CC BY 4.0, interactive
- Dive into Deep Learning, ch. 12: Optimisation โ ยท Zhang, Lipton, Li & Smola ยท CC BY-SA 4.0