OpenAI's GPT-1: generative pre-training of Transformers
OpenAI showed that pre-training a Transformer language model on unlabeled text and then fine-tuning it yields strong results across many NLP tasks — the first 'GPT'.
Key facts
- Paper: 'Improving Language Understanding by Generative Pre-Training' (Radford et al.)
- ~117M parameters, 12-layer decoder-only Transformer
- Pre-trained on the BooksCorpus dataset
- Improved state of the art on 9 of 12 benchmarks studied
What happened
OpenAI published a semi-supervised approach: unsupervised generative pre-training followed by supervised fine-tuning.
Why it matters
Established the pre-train-then-adapt paradigm and the decoder-only Transformer lineage that led to ChatGPT.
Changelog
- 2026-09-29: created
Related events
- 'Attention Is All You Need' introduces the Transformer ★★★★★
- OpenAI announces GPT-2 and withholds the full model over misuse concerns ★★★★
- Google releases BERT, bidirectional Transformer pre-training ★★★★
Sources (2)
id: 2018-06-11-gpt-1 · updated 2026-09-29 · open in the interactive timeline