Backpropagation: blame, passed backwards

How a network with millions of weights figures out which ones to adjust.

⏱ 6 min read

Gradient descent needs the slope for every weight. With millions of weights, how do we get them all efficiently? Backpropagation: one pass forward to make a prediction, one pass backward to share out the blame.

forward: make a predictionLosshow wrong?backward: share out the blame (gradients)inputhiddenoutput
Forward pass computes the prediction and loss. Backward pass sends error signals back through each layer to get every weight's gradient.

➑️ Forward pass

Input flows through the layers, each doing its weighted sums and activations, until we get a prediction.

Compare to the truth β†’ loss.

⬅️ Backward pass

Starting from the loss, ask each layer: "how much did you contribute to this error?"

The chain rule from calculus lets us compute this layer by layer, reusing work from the layer above. That reuse is why it's fast.

🧩 Quick quiz

What mathematical rule makes backpropagation work?

πŸ«₯ Vanishing & exploding gradients

Multiplying many small numbers through deep networks can shrink gradients to almost zero (vanishing), so early layers stop learning. Multiplying big ones can blow up (exploding).

Fixes: ReLU activations, careful initialisation, normalisation layers, and residual connections (shortcuts that let gradients skip layers).

πŸ€– You rarely write it by hand

Frameworks like PyTorch and JAX do automatic differentiation: you write the forward pass and they compute all the gradients for you.

Still, building it once yourself (see Karpathy's micrograd below) makes everything click.

🧩 Quick quiz

Residual (skip) connections mainly help with…

✨ Before you drift off

  • Forward pass β†’ prediction & loss; backward pass β†’ gradients.
  • The chain rule spreads blame layer by layer efficiently.
  • Residual connections and normalisation tame vanishing gradients.

πŸ“š Go deeper (free & open)