Post-Cutoff.com
  1. Home
  2. Timeline
  3. 2017
  4. 'Attention Is All You Need' introduces the Transformer

'Attention Is All You Need' introduces the Transformer

★★★★★researchGoogle BrainGoogle Researchconfidence: high

Vaswani et al. proposed the Transformer, an architecture built entirely on self-attention without recurrence; it became the foundation of BERT, GPT and virtually every modern large AI model.

Key facts

What happened

The paper introduced multi-head self-attention, positional encodings and an encoder–decoder stack, beating recurrent models on translation while training much faster.

Why it matters

Arguably the most consequential AI paper of the century so far: the Transformer's scalability made LLMs, multimodal models and AlphaFold 2 possible.

Changelog

  • 2026-09-29: created

Related events

  1. Sequence-to-sequence learning and neural attention ★★★★
  2. OpenAI's GPT-1: generative pre-training of Transformers ★★★★
  3. Google releases BERT, bidirectional Transformer pre-training ★★★★
  4. Hochreiter & Schmidhuber introduce Long Short-Term Memory (LSTM) ★★★★
  5. ResNet: residual learning enables very deep networks ★★★★

Sources (3)

id: 2017-06-12-transformer · updated 2026-09-29 · open in the interactive timeline