Scaling laws, evals & safety
Why bigger got better, how we measure frontier models, and the open problems in keeping AI safe.
Much of the last decade of AI progress came from one discovery: performance improves predictably as you scale up compute, data and model size. That predictability shapes billion-dollar decisions, and raises serious questions.
๐ Scaling laws
Kaplan et al. (2020) found loss follows smooth power laws in parameters, data and compute.
Chinchilla (2022) showed many models were under-trained: for a fixed compute budget, scale parameters and tokens together (roughly 20 tokens per parameter).
What was the key Chinchilla finding?
๐งช Evaluating frontier models
Benchmarks test knowledge (MMLU), maths (GSM8K, MATH), coding (HumanEval, SWE-bench) and more.
Problems: benchmarks saturate, and test questions can leak into training data (contamination). So labs increasingly use held-out, human-rated and agentic evaluations.
๐ก๏ธ Alignment & safety
Alignment: making systems pursue what we actually intend. Reward hacking and sycophancy are everyday examples of misalignment.
Robustness: resisting jailbreaks and prompt injection.
Interpretability: understanding what's happening inside the network, e.g. finding human-understandable features inside models.
๐๏ธ Responsible deployment
Labs publish model cards and system cards, run red-teaming, and test for dangerous capabilities before release.
Policies like responsible scaling commitments tie stronger safeguards to more capable models.
What is "benchmark contamination"?
โจ Before you drift off
- Loss falls predictably with scale (power laws).
- Compute-optimal training scales data with parameters.
- Benchmarks saturate and leak, so evaluation is an active field.
- Alignment, robustness and interpretability are core open problems.
๐ Go deeper (free & open)
- Scaling Laws for Neural Language Models โ ยท Kaplan et al., 2020
- Training Compute-Optimal LLMs (Chinchilla) โ ยท Hoffmann et al., 2022
- Concrete Problems in AI Safety โ ยท Amodei et al., 2016
- Transformer Circuits Thread (interpretability research) โ ยท Anthropic interpretability team