DeepSeek-V3: frontier-level open model trained for ~$5.6M in GPU time
Chinese lab DeepSeek released DeepSeek-V3, a 671B-parameter mixture-of-experts model (37B active) with open weights that rivaled GPT-4o and Claude 3.5 Sonnet; its final training run reportedly used 2.788M H800 GPU-hours (~$5.6M).
Key facts
- Released 26 December 2024; technical report arXiv 2412.19437
- 671B total parameters, 37B activated per token
- Pre-trained on 14.8 trillion tokens
- 2.788M H800 GPU-hours for full training (~$5.576M at $2/GPU-hour, excluding prior research)
- Innovations: multi-head latent attention, auxiliary-loss-free load balancing, FP8 training, multi-token prediction
What happened
DeepSeek published open weights and an unusually detailed report showing frontier performance at a fraction of the reported compute of US labs, despite export controls.
Why it matters
Upended assumptions about the cost of frontier AI and China's position; it was the base for DeepSeek-R1 weeks later.
Changelog
- 2026-09-29: created
Related events
- DeepSeek-R1: open-weights reasoning model rivals o1 and shakes markets ★★★★★
- Mistral AI releases Mixtral 8x7B, an open mixture-of-experts model ★★★
- DeepSeekMath-V2: open-weights self-verifying prover reaches IMO 2025 gold level and 118/120 on Putnam 2024 ★★★★
Sources (2)
id: 2024-12-26-deepseek-v3 · updated 2026-09-29 · open in the interactive timeline