Gradient descent: rolling downhill

How models find good weights: feel the slope, take a small step down, repeat.

โฑ 6 min read

Picture yourself on a foggy hillside at night. You can't see the valley, but you can feel which way the ground slopes under your feet. Step downhill. Feel again. Step again. That's gradient descent, the algorithm that trains almost every model in existence.

loss โ†‘weight โ†’startโ† minimum โœ“steep slope โ†’ big stepsgentle slope โ†’ tiny steps
The loss curve is the hill. Each step moves the weight a little in the downhill direction until we settle near the bottom.

๐Ÿ“ The gradient

The gradient is the slope of the loss with respect to each weight. It points uphill.

So we move in the opposite direction: w_new = w_old โˆ’ learning_rate ร— gradient.

๐Ÿ‘Ÿ The learning rate

The learning rate is the step size.

Too small: you'll get there, eventually, in a few weeks. Too big: you leap over the valley and bounce around or fly off entirely.

Picking it well is one of the most important practical skills in deep learning.

๐Ÿงฉ Quick quiz

Training loss is jumping up and down wildly and getting worse. What's the most likely fix?

๐ŸŽ’ Mini-batches (SGD)

Computing the gradient on the whole dataset each step is slow. Instead we use a small random mini-batch (say 32 or 256 examples).

This is stochastic gradient descent (SGD). The noisy steps are faster, and the noise even helps escape bad spots.

๐Ÿš€ Fancier optimisers

Momentum keeps rolling in a consistent direction, like a ball gathering speed.

Adam adapts the step size for every weight individually. It's the default choice for most deep learning today.

๐Ÿงฉ Quick quiz

Which way does gradient descent step?

โœจ Before you drift off

  • Gradient = slope of the loss; it points uphill.
  • Update: w โ† w โˆ’ learning_rate ร— gradient.
  • Learning rate too high โ†’ bouncing; too low โ†’ crawling.
  • Mini-batches (SGD) and Adam make it fast in practice.

๐Ÿ“š Go deeper (free & open)