Overfitting & the Goldilocks model

Too simple misses the pattern, too complex memorises the noise. How to land just right.

⏱ 5 min read

A student who memorises every past exam word for word might ace those papers and then flunk the real exam. Models do the same thing. It's called overfitting, and fighting it is a huge part of machine learning.

Underfit 😴Just right 😊Overfit πŸ€“
Underfit (too simple), just right, and overfit (wiggles through every point, including the noise).

😴 Underfitting

The model is too simple to capture the pattern. It does badly on training data and new data.

Fix: a bigger model, better features, or train longer.

πŸ€“ Overfitting

The model memorises the training data, noise and all. Training error is tiny but validation error is high.

Signature: the training loss keeps falling while the validation loss starts rising.

🧩 Quick quiz

Training accuracy: 99%. Validation accuracy: 71%. What's going on?

πŸ›‘οΈ Ways to fight overfitting

More data: the best cure when you can get it.

Data augmentation: flip, crop, rotate images to make "new" examples.

Regularisation (L1/L2): penalise large weights so the model stays simple.

Dropout: randomly switch off neurons during training so none can be relied on too much.

Early stopping: stop training when validation loss stops improving.

βš–οΈ The bias–variance trade-off

Bias: error from wrong assumptions (too simple β†’ underfit).

Variance: error from being too sensitive to the specific training data (too complex β†’ overfit).

Good models balance the two.

🧩 Quick quiz

Which of these is NOT a way to reduce overfitting?

✨ Before you drift off

  • Underfit = too simple; overfit = memorised the noise.
  • Watch for the gap between training and validation scores.
  • Fight overfitting with more data, augmentation, regularisation, dropout, early stopping.

πŸ“š Go deeper (free & open)