OpenAI unveils DALL·E and CLIP
DALL·E generated images from text prompts using a 12B-parameter Transformer, and CLIP learned joint image–text representations from 400M image-caption pairs; CLIP became a key component of later diffusion image generators.
Key facts
- Both announced 5 January 2021
- DALL·E: 12-billion-parameter version of GPT-3 trained on text–image pairs
- CLIP: trained on 400M image–text pairs; strong zero-shot ImageNet accuracy
- CLIP weights open-sourced; used by Stable Diffusion's text encoder (v1)
What happened
OpenAI introduced a text-to-image model and a contrastive vision-language model on the same day.
Why it matters
Launched the text-to-image era and made natural language the interface for vision models.
Changelog
- 2026-09-29: created
Related events
- DALL·E 2 brings photorealistic text-to-image generation ★★★★
- Stable Diffusion released as open weights ★★★★★
Sources (4)
- officialDALL·E: Creating images from text (OpenAI)
- officialCLIP: Connecting text and images (OpenAI)
- paperLearning Transferable Visual Models From Natural Language Supervision (arXiv)
- paperZero-Shot Text-to-Image Generation (arXiv)
id: 2021-01-05-dall-e-clip · updated 2026-09-29 · open in the interactive timeline