DeepSeek V4 preview: 1.6T-parameter open MoE running on Huawei Ascend
DeepSeek released a preview of V4 on 2026-04-24: V4-Pro (1.6T total / 49B active) and V4-Flash (284B / 13B active), both MIT-licensed MoE models with a 1M-token context, validated on Huawei Ascend NPUs as well as Nvidia GPUs, priced far below Western frontier APIs.
Key facts
- V4-Pro: 1.6T total parameters, 49B active; V4-Flash: 284B total, 13B active (The Register)
- Training data: 33T tokens; context window 1M tokens
- KV cache 9.5x-13.7x smaller than DeepSeek V3.2; mixed FP8/FP4 precision with quantization-aware training of MoE experts
- New hybrid attention (Compressed Sparse Attention + Heavily Compressed Attention) and Muon optimizer
- API price: Flash $0.14/M input, $0.28/M output; Pro $1.74/M input, $3.48/M output
- Day-zero support on Huawei Ascend SuperNode line incl. Ascend 950; weights on Hugging Face under MIT license
What happened
On Friday 2026-04-24 DeepSeek published a preview of its fourth-generation model family. Two MoE models shipped: V4-Pro (1.6 trillion parameters, 49B active) and V4-Flash (284B, 13B active), both with a 1M-token context window and trained on ~33T tokens. Architecturally, DeepSeek introduced a hybrid compressed attention scheme and adopted the Muon optimizer, and cut KV-cache memory 9.5-13.7x versus V3.2, using FP8/FP4 mixed precision with quantization-aware training.
The launch was notable for hardware: DeepSeek validated the models on Huawei Ascend NPUs (Huawei announced day-zero support across its SuperNode line, including Ascend 950) as well as Nvidia GPUs. Coverage (Tom's Hardware) linked the release to escalating US government accusations of IP theft / distillation by Chinese labs. Later milestones: V4-Flash re-post-trained update (2026-07-31), V4-Pro GA with low/high/max thinking effort (2026-08-13), and V4.1-Flash (2026-09-10).
Why it matters
V4 was the largest open-weights model at release and the first frontier-class release optimized for a Chinese AI accelerator, a signal that China's model stack can decouple from Nvidia. Its aggressive pricing (Pro output $3.48/M) kept pressure on Western API prices.
Changelog
- 2026-09-29: created
Related events
- DeepSeek V4.1-Flash: new architecture family, native vision, cheaper API ★★★
- Huawei sets Ascend 950 cluster cloud launch (China Sept 30, global Nov 30) and Ascend 960 roadmap ★★★
Sources (4)
- pressThe Register: DeepSeek's new models offer big inference cost savings
- pressTom's Hardware: DeepSeek launches 1.6T V4 on Huawei chips
- pressHuawei Central: DeepSeek launches V4 on Huawei chips
- docsDeepSeek API changelog
id: 2026-04-24-deepseek-v4-preview · updated 2026-09-29 · open in the interactive timeline