Is my model any good?
Accuracy can lie. Meet the confusion matrix, precision, recall and F1.
A model that predicts "no fraud" for every transaction is 99.9% accurate if only 0.1% of transactions are fraud, and completely useless. To judge classifiers properly, you need better tools than accuracy.
π² The four boxes
True Positive (TP): said yes, was yes. π
True Negative (TN): said no, was no. π
False Positive (FP): said yes, was actually no (false alarm).
False Negative (FN): said no, was actually yes (missed it).
π― Precision
Precision = TP / (TP + FP). Of everything I flagged, how much was right?
Matters when false alarms are costly, e.g. a spam filter binning your job offer.
πΈοΈ Recall
Recall = TP / (TP + FN). Of all the real positives, how many did I catch?
Matters when misses are costly, e.g. screening for cancer.
For an airport weapon scanner, which metric matters most?
βοΈ F1 score & the trade-off
Raising your threshold usually boosts precision but lowers recall, and vice versa.
F1 is the harmonic mean of the two, a single number that's only high when both are.
π For regression
MAE: average absolute error, easy to read ("off by Β£12k on average").
RMSE: like MAE but punishes big misses more.
RΒ²: how much of the variation your model explains (1.0 is perfect).
Your model flagged 100 emails as spam; 90 really were spam. What's the precision?
β¨ Before you drift off
- Accuracy misleads on imbalanced data.
- Precision: how many flagged were right. Recall: how many real ones were caught.
- F1 balances both; pick metrics based on the cost of each mistake.
π Go deeper (free & open)
- ML Crash Course: Classification metrics β Β· Google Β· CC BY 4.0
- scikit-learn: Model evaluation guide β Β· scikit-learn developers Β· BSD-3