Google releases BERT, bidirectional Transformer pre-training
BERT pre-trained a bidirectional Transformer encoder with masked language modeling and set new records on 11 NLP tasks; it was open-sourced and soon deployed in Google Search.
Key facts
- arXiv 1810.04805 (October 2018); NAACL 2019 best paper
- BERT-Large: 340M parameters
- Pre-training objectives: masked LM + next sentence prediction
- Google said in October 2019 that BERT was used in Search ranking
What happened
Google published and open-sourced BERT, which fine-tuned easily to classification, QA and tagging tasks.
Why it matters
Triggered the 'ImageNet moment' of NLP: pre-trained Transformers became the default for all language tasks.
Changelog
- 2026-09-29: created
Related events
- 'Attention Is All You Need' introduces the Transformer ★★★★★
- OpenAI's GPT-1: generative pre-training of Transformers ★★★★
Sources (2)
id: 2018-10-11-bert · updated 2026-09-29 · open in the interactive timeline