Post-Cutoff.com
  1. Home
  2. Timeline
  3. 2017
  4. AlphaGo Zero and AlphaZero master games through pure…

AlphaGo Zero and AlphaZero master games through pure self-play

★★★★★researchDeepMindconfidence: high

AlphaGo Zero (Nature, October 2017) learned Go from scratch with no human games and beat the version that defeated Lee Sedol 100–0; AlphaZero (December 2017) generalized the method to chess and shogi.

Key facts

What happened

DeepMind showed that a single algorithm combining a neural network with tree search, trained only by playing against itself, reached superhuman strength in three classic board games.

Why it matters

Proved that learning from self-generated experience can exceed human knowledge — an idea that resurfaced in RL-trained reasoning models (o1, R1) in 2024–2025.

Changelog

  • 2026-09-29: created

Related events

  1. AlphaGo defeats Lee Sedol 4–1 at Go ★★★★★
  2. OpenAI o1: reasoning models trained with reinforcement learning ★★★★★

Sources (3)

id: 2017-12-05-alphazero · updated 2026-09-29 · open in the interactive timeline