DeepMind's Chinchilla revises scaling laws toward more data
Hoffmann et al. found that for compute-optimal training, parameters and training tokens should scale equally (~20 tokens per parameter); 70B Chinchilla outperformed the 280B Gopher.
Key facts
- arXiv 2203.15556 'Training Compute-Optimal Large Language Models'
- Chinchilla: 70B parameters trained on 1.4 trillion tokens
- Beat Gopher (280B), GPT-3 (175B) and Megatron-Turing NLG (530B) on many benchmarks
- Implied most prior LLMs were undertrained
What happened
Over 400 training runs showed earlier scaling laws had over-weighted parameter count relative to data.
Why it matters
Reshaped how every lab trains LLMs, pushing toward far larger datasets and smaller, cheaper-to-serve models (e.g. LLaMA).
Changelog
- 2026-09-29: created
Related events
- OpenAI publishes 'Scaling Laws for Neural Language Models' ★★★★★
- Meta releases LLaMA, sparking the open-weights LLM wave ★★★★★
Sources (2)
- paperTraining Compute-Optimal Large Language Models (arXiv)
- discussionWikipedia: Chinchilla (language model)
id: 2022-03-29-chinchilla · updated 2026-09-29 · open in the interactive timeline