Backpropagation: blame, passed backwards
How a network with millions of weights figures out which ones to adjust.
Gradient descent needs the slope for every weight. With millions of weights, how do we get them all efficiently? Backpropagation: one pass forward to make a prediction, one pass backward to share out the blame.
β‘οΈ Forward pass
Input flows through the layers, each doing its weighted sums and activations, until we get a prediction.
Compare to the truth β loss.
β¬ οΈ Backward pass
Starting from the loss, ask each layer: "how much did you contribute to this error?"
The chain rule from calculus lets us compute this layer by layer, reusing work from the layer above. That reuse is why it's fast.
What mathematical rule makes backpropagation work?
π«₯ Vanishing & exploding gradients
Multiplying many small numbers through deep networks can shrink gradients to almost zero (vanishing), so early layers stop learning. Multiplying big ones can blow up (exploding).
Fixes: ReLU activations, careful initialisation, normalisation layers, and residual connections (shortcuts that let gradients skip layers).
π€ You rarely write it by hand
Frameworks like PyTorch and JAX do automatic differentiation: you write the forward pass and they compute all the gradients for you.
Still, building it once yourself (see Karpathy's micrograd below) makes everything click.
Residual (skip) connections mainly help withβ¦
β¨ Before you drift off
- Forward pass β prediction & loss; backward pass β gradients.
- The chain rule spreads blame layer by layer efficiently.
- Residual connections and normalisation tame vanishing gradients.
π Go deeper (free & open)
- Neural Networks: Zero to Hero (micrograd lecture) β Β· Andrej Karpathy Β· Build backprop from scratch, MIT-licensed code
- What is backpropagation really doing? β Β· 3Blue1Brown
- CS231n: Backpropagation notes β Β· Stanford University