Is my model any good?

Accuracy can lie. Meet the confusion matrix, precision, recall and F1.

⏱ 6 min read

A model that predicts "no fraud" for every transaction is 99.9% accurate if only 0.1% of transactions are fraud, and completely useless. To judge classifiers properly, you need better tools than accuracy.

MODEL PREDICTEDYesNoACTUALLYYesNoTPcaught it βœ“FNmissed it βœ—FPfalse alarm βœ—TNcorrectly ignored βœ“
The confusion matrix: every prediction lands in one of four boxes.

πŸ”² The four boxes

True Positive (TP): said yes, was yes. πŸŽ‰

True Negative (TN): said no, was no. πŸ‘

False Positive (FP): said yes, was actually no (false alarm).

False Negative (FN): said no, was actually yes (missed it).

🎯 Precision

Precision = TP / (TP + FP). Of everything I flagged, how much was right?

Matters when false alarms are costly, e.g. a spam filter binning your job offer.

πŸ•ΈοΈ Recall

Recall = TP / (TP + FN). Of all the real positives, how many did I catch?

Matters when misses are costly, e.g. screening for cancer.

🧩 Quick quiz

For an airport weapon scanner, which metric matters most?

βš–οΈ F1 score & the trade-off

Raising your threshold usually boosts precision but lowers recall, and vice versa.

F1 is the harmonic mean of the two, a single number that's only high when both are.

πŸ“ˆ For regression

MAE: average absolute error, easy to read ("off by Β£12k on average").

RMSE: like MAE but punishes big misses more.

RΒ²: how much of the variation your model explains (1.0 is perfect).

🧩 Quick quiz

Your model flagged 100 emails as spam; 90 really were spam. What's the precision?

✨ Before you drift off

  • Accuracy misleads on imbalanced data.
  • Precision: how many flagged were right. Recall: how many real ones were caught.
  • F1 balances both; pick metrics based on the cost of each mistake.

πŸ“š Go deeper (free & open)