OpenAI publishes 'Scaling Laws for Neural Language Models'
Kaplan et al. showed language-model loss falls as a smooth power law in parameters, data and compute over many orders of magnitude, giving a quantitative case for building ever-larger models.
Key facts
- arXiv 2001.08361 (January 2020)
- Loss follows power laws in model size, dataset size and compute
- Architecture details (depth/width) matter far less than scale
- Later revised by DeepMind's Chinchilla (2022) on the optimal data/parameter ratio
What happened
The paper fit empirical power laws across hundreds of training runs and derived compute-optimal allocation rules.
Why it matters
Scaling laws became the strategic basis for the trillion-dollar compute build-out of the 2020s.
Changelog
- 2026-09-29: created
Related events
- GPT-3 (175B) shows in-context few-shot learning ★★★★★
- DeepMind's Chinchilla revises scaling laws toward more data ★★★★
- Rich Sutton publishes "The Bitter Lesson": general methods that scale with compute win ★★★★
Sources (2)
id: 2020-01-23-scaling-laws · updated 2026-09-29 · open in the interactive timeline