Data is everything
A model is only as good as what it learned from. Splits, bias, and why "garbage in, garbage out" still rules.
Engineers joke that machine learning is 80% data cleaning and 20% complaining about data cleaning. It's funny because it's true: the data shapes everything a model will ever know.
๐งช Never grade with the practice questions
If a student sees the exam answers while studying, a perfect score proves nothing.
So we split data: most for training, some held back to check progress (validation), and some locked away for a final honest exam (test).
Why do we keep a separate test set?
๐๏ธ Garbage in, garbage out
Wrong labels, duplicates, and missing values all teach the model wrong lessons.
A smaller clean dataset often beats a bigger messy one.
โ๏ธ Bias in, bias out
If a hiring model learns from past decisions that were unfair, it learns to be unfair too, and does it at scale.
If a face-recognition dataset is mostly one skin tone, it will work worse for everyone else. Real systems have failed exactly this way.
๐ข Features
The pieces of information we feed a model are called features: for a house, things like size, location and age.
Deep learning can learn its own features from raw data (pixels, words), which is a big reason it took off.
A loan model was trained only on applicants from big cities. What's the most likely problem?
โจ Before you drift off
- Split data into train / validation / test, and keep the test set untouched.
- Data quality usually matters more than data quantity.
- Biased data produces biased models; ask who is missing.
๐ Go deeper (free & open)
- ML Crash Course: Working with data & Fairness โ ยท Google ยท CC BY 4.0
- Datasheets for Datasets โ ยท Gebru et al. ยท Research paper on documenting datasets
- Ethics of AI (free course) โ ยท University of Helsinki