Sequence-to-sequence learning and neural attention
Sutskever, Vinyals and Le's seq2seq (LSTM encoder–decoder) and Bahdanau, Cho and Bengio's attention mechanism, both posted in September 2014, made end-to-end neural machine translation work.
Key facts
- Bahdanau et al. attention paper: arXiv 1409.0473 (1 Sep 2014)
- Sutskever et al. seq2seq paper: arXiv 1409.3215 (10 Sep 2014)
- Google Neural Machine Translation system launched in 2016 (arXiv 1609.08144)
- Attention later became the sole core mechanism of the Transformer
What happened
Two papers showed neural networks could map whole sequences to sequences, and that letting the decoder 'attend' to encoder states greatly improved long sentences.
Why it matters
Established the encoder–decoder paradigm and attention — the direct precursors of the Transformer and modern LLMs.
Changelog
- 2026-09-29: created
Related events
- Hochreiter & Schmidhuber introduce Long Short-Term Memory (LSTM) ★★★★
- 'Attention Is All You Need' introduces the Transformer ★★★★★
Sources (3)
- paperSequence to Sequence Learning with Neural Networks (arXiv)
- paperNeural Machine Translation by Jointly Learning to Align and Translate (arXiv)
- paperGoogle's Neural Machine Translation System (arXiv)
id: 2014-09-10-seq2seq-attention · updated 2026-09-29 · open in the interactive timeline