Scaling laws, evals & safety

Why bigger got better, how we measure frontier models, and the open problems in keeping AI safe.

โฑ 7 min read

Much of the last decade of AI progress came from one discovery: performance improves predictably as you scale up compute, data and model size. That predictability shapes billion-dollar decisions, and raises serious questions.

10^1810^2010^2210^24training compute (FLOPs, log scale) โ†’loss (log) โ†“ betterpredictedeach dot = one model size
On a logโ€“log plot, loss falls in a remarkably straight line as training compute grows (illustrative).

๐Ÿ“ˆ Scaling laws

Kaplan et al. (2020) found loss follows smooth power laws in parameters, data and compute.

Chinchilla (2022) showed many models were under-trained: for a fixed compute budget, scale parameters and tokens together (roughly 20 tokens per parameter).

๐Ÿงฉ Quick quiz

What was the key Chinchilla finding?

๐Ÿงช Evaluating frontier models

Benchmarks test knowledge (MMLU), maths (GSM8K, MATH), coding (HumanEval, SWE-bench) and more.

Problems: benchmarks saturate, and test questions can leak into training data (contamination). So labs increasingly use held-out, human-rated and agentic evaluations.

๐Ÿ›ก๏ธ Alignment & safety

Alignment: making systems pursue what we actually intend. Reward hacking and sycophancy are everyday examples of misalignment.

Robustness: resisting jailbreaks and prompt injection.

Interpretability: understanding what's happening inside the network, e.g. finding human-understandable features inside models.

๐Ÿ›๏ธ Responsible deployment

Labs publish model cards and system cards, run red-teaming, and test for dangerous capabilities before release.

Policies like responsible scaling commitments tie stronger safeguards to more capable models.

๐Ÿงฉ Quick quiz

What is "benchmark contamination"?

โœจ Before you drift off

  • Loss falls predictably with scale (power laws).
  • Compute-optimal training scales data with parameters.
  • Benchmarks saturate and leak, so evaluation is an active field.
  • Alignment, robustness and interpretability are core open problems.

๐Ÿ“š Go deeper (free & open)