Data is everything

A model is only as good as what it learned from. Splits, bias, and why "garbage in, garbage out" still rules.

โฑ 5 min read

Engineers joke that machine learning is 80% data cleaning and 20% complaining about data cleaning. It's funny because it's true: the data shapes everything a model will ever know.

๐Ÿงช Never grade with the practice questions

If a student sees the exam answers while studying, a perfect score proves nothing.

So we split data: most for training, some held back to check progress (validation), and some locked away for a final honest exam (test).

Train70%learn from itVal15%tuneTest15%final ๐Ÿ”’
A common split. The test set is only opened once, at the very end.
๐Ÿงฉ Quick quiz

Why do we keep a separate test set?

๐Ÿ—‘๏ธ Garbage in, garbage out

Wrong labels, duplicates, and missing values all teach the model wrong lessons.

A smaller clean dataset often beats a bigger messy one.

โš–๏ธ Bias in, bias out

If a hiring model learns from past decisions that were unfair, it learns to be unfair too, and does it at scale.

If a face-recognition dataset is mostly one skin tone, it will work worse for everyone else. Real systems have failed exactly this way.

๐Ÿ”ข Features

The pieces of information we feed a model are called features: for a house, things like size, location and age.

Deep learning can learn its own features from raw data (pixels, words), which is a big reason it took off.

๐Ÿงฉ Quick quiz

A loan model was trained only on applicants from big cities. What's the most likely problem?

โœจ Before you drift off

  • Split data into train / validation / test, and keep the test set untouched.
  • Data quality usually matters more than data quantity.
  • Biased data produces biased models; ask who is missing.

๐Ÿ“š Go deeper (free & open)