'Attention Is All You Need' introduces the Transformer
Vaswani et al. proposed the Transformer, an architecture built entirely on self-attention without recurrence; it became the foundation of BERT, GPT and virtually every modern large AI model.
Key facts
- arXiv 1706.03762, posted 12 June 2017; NeurIPS 2017
- Eight co-authors from Google Brain/Research
- WMT 2014 English–German: 28.4 BLEU, a new state of the art
- Highly parallelizable training vs. RNNs, enabling scale
- The 'T' in GPT stands for Transformer
What happened
The paper introduced multi-head self-attention, positional encodings and an encoder–decoder stack, beating recurrent models on translation while training much faster.
Why it matters
Arguably the most consequential AI paper of the century so far: the Transformer's scalability made LLMs, multimodal models and AlphaFold 2 possible.
Changelog
- 2026-09-29: created
Related events
- Sequence-to-sequence learning and neural attention ★★★★
- OpenAI's GPT-1: generative pre-training of Transformers ★★★★
- Google releases BERT, bidirectional Transformer pre-training ★★★★
- Hochreiter & Schmidhuber introduce Long Short-Term Memory (LSTM) ★★★★
- ResNet: residual learning enables very deep networks ★★★★
Sources (3)
- paperAttention Is All You Need (arXiv)
- officialGoogle Research blog: Transformer
- discussionWikipedia: Attention Is All You Need
id: 2017-06-12-transformer · updated 2026-09-29 · open in the interactive timeline