Overfitting & the Goldilocks model
Too simple misses the pattern, too complex memorises the noise. How to land just right.
A student who memorises every past exam word for word might ace those papers and then flunk the real exam. Models do the same thing. It's called overfitting, and fighting it is a huge part of machine learning.
π΄ Underfitting
The model is too simple to capture the pattern. It does badly on training data and new data.
Fix: a bigger model, better features, or train longer.
π€ Overfitting
The model memorises the training data, noise and all. Training error is tiny but validation error is high.
Signature: the training loss keeps falling while the validation loss starts rising.
Training accuracy: 99%. Validation accuracy: 71%. What's going on?
π‘οΈ Ways to fight overfitting
More data: the best cure when you can get it.
Data augmentation: flip, crop, rotate images to make "new" examples.
Regularisation (L1/L2): penalise large weights so the model stays simple.
Dropout: randomly switch off neurons during training so none can be relied on too much.
Early stopping: stop training when validation loss stops improving.
βοΈ The biasβvariance trade-off
Bias: error from wrong assumptions (too simple β underfit).
Variance: error from being too sensitive to the specific training data (too complex β overfit).
Good models balance the two.
Which of these is NOT a way to reduce overfitting?
β¨ Before you drift off
- Underfit = too simple; overfit = memorised the noise.
- Watch for the gap between training and validation scores.
- Fight overfitting with more data, augmentation, regularisation, dropout, early stopping.
π Go deeper (free & open)
- ML Crash Course: Overfitting & regularisation β Β· Google Β· CC BY 4.0
- scikit-learn: Underfitting vs. Overfitting example β Β· scikit-learn developers Β· BSD-3
- Dropout paper (JMLR 2014) β Β· Srivastava, Hinton et al.