AlphaGo Zero and AlphaZero master games through pure self-play
AlphaGo Zero (Nature, October 2017) learned Go from scratch with no human games and beat the version that defeated Lee Sedol 100–0; AlphaZero (December 2017) generalized the method to chess and shogi.
Key facts
- AlphaGo Zero Nature paper published 18 October 2017
- AlphaGo Zero beat AlphaGo Lee 100–0
- AlphaZero preprint arXiv 1712.01815 (5 December 2017); Science paper December 2018
- AlphaZero defeated Stockfish (chess) and Elmo (shogi) after hours of self-play training
What happened
DeepMind showed that a single algorithm combining a neural network with tree search, trained only by playing against itself, reached superhuman strength in three classic board games.
Why it matters
Proved that learning from self-generated experience can exceed human knowledge — an idea that resurfaced in RL-trained reasoning models (o1, R1) in 2024–2025.
Changelog
- 2026-09-29: created
Related events
- AlphaGo defeats Lee Sedol 4–1 at Go ★★★★★
- OpenAI o1: reasoning models trained with reinforcement learning ★★★★★
Sources (3)
- paperMastering Chess and Shogi by Self-Play with a General RL Algorithm (arXiv)
- paperMastering the game of Go without human knowledge (Nature, DOI)
- paperA general reinforcement learning algorithm that masters chess, shogi, and Go (Science, DOI)
id: 2017-12-05-alphazero · updated 2026-09-29 · open in the interactive timeline