Lines and curves: regression
The humble straight line is the ancestor of every neural network. Linear and logistic regression, explained.
Before neural networks, there was the line. Understanding the two classic models, linear regression and logistic regression, gives you 80% of the intuition for everything that follows.
📏 Linear regression
Predict a number from inputs with y = w·x + b. With many inputs: y = w₁x₁ + w₂x₂ + … + b.
Example: house price = 3,000 × size + 20,000 × bedrooms + 50,000.
"Training" means finding the ws and b that make the line fit the data best.
📐 Measuring fit: Mean Squared Error
For each point, take the gap between prediction and truth, square it (so negatives don't cancel and big misses hurt more), then average.
MSE = average of (prediction − truth)². Smaller is better.
Why does MSE square the errors instead of just averaging them?
🔀 Logistic regression (it's for classification!)
Despite the name, it predicts yes/no. Take the linear score and pass it through the sigmoid function σ(z) = 1 / (1 + e⁻ᶻ).
Sigmoid squashes any number into 0–1, which we read as a probability: 0.92 → "92% likely spam".
🔗 The link to neural networks
A single neuron with a sigmoid activation is logistic regression.
A neural network is many of these stacked and connected. So you already understand the building block.
Logistic regression outputs…
✨ Before you drift off
- Linear regression predicts numbers: y = w·x + b.
- MSE measures fit by averaging squared errors.
- Logistic regression = linear score + sigmoid → probability.
- One sigmoid neuron is logistic regression.
📚 Go deeper (free & open)
- ML Crash Course: Linear & logistic regression ↗ · Google · CC BY 4.0
- CS229 Lecture Notes ↗ · Stanford University (Andrew Ng et al.)
- Dive into Deep Learning, ch. 3–4 ↗ · Zhang, Lipton, Li & Smola · CC BY-SA 4.0, with runnable code