DeepSeek-V4.1-Flash
Call as deepseek-flash. Legacy ids deepseek-v4-flash and deepseek-v4-flash-vision-exp are routed here and billed at Flash price. Knowledge cutoff not published.
- Context window
- 1,000,000 tokens
- Max output
- 384,000 tokens
- Input
- text, image
- Output
- text
- License
- mit
- Pricing
- input: $0.3 · output: $1.2 · cache read: $0.006 (per 1M tokens (USD), peak-hour list price; off-peak is half (input 0.15, output 0.6, cache hit 0.003). Peak = 01:00-04:00 and 06:00-10:00 UTC Mon-Fri) source
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| DeepSeek API | deepseek-flash | https://api.deepseek.com | docs |
| DeepSeek API (Anthropic format) | deepseek-flash | https://api.deepseek.com/anthropic | docs |
| Alibaba Cloud Model Studio | deepseek-v4.1-flash | — | docs |
| OpenRouter | deepseek/deepseek-v4.1-flash | openrouter.ai/deepseek/deepseek-v4.1-flash | — |
| Hugging Face | — | huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash | — |
| Web app | — | chat.deepseek.com | — |
Notable capabilities (5)
- Native vision in the Flash tier: First DeepSeek Flash model with native multimodal (image) understanding built in; replaced the separate V4-Flash-Vision-Exp. source
- Causal Encoder-Decoder (CED) architecture: 552B-backbone MoE that activates only ~8B params per token in prefill and ~16B in decode, aimed at input-heavy agentic workloads. source
- Tiny KV cache (CSA2 + FP4 KV): Compressed Sparse Attention 2 and FP4 main KV cache cut the global KV cache to ~890 bytes/token, about 1/4 of V4-Flash. source
- Hybrid thinking with effort levels: One model id serves thinking (default) and non-thinking modes; reasoning effort low/high/max. source
- Multiple API protocols: Same model served via OpenAI Chat Completions, OpenAI Responses (Codex-adapted) and Anthropic Messages formats. source
DeepSeek's cheapest current model: agentic coding, long-context (1M) work and image understanding at very low cost. Schedule batch jobs off-peak for 50% off.
curl https://api.deepseek.com/chat/completions \
-H "Authorization: Bearer $DEEPSEEK_API_KEY" -H "Content-Type: application/json" \
-d '{"model":"deepseek-flash","messages":[{"role":"user","content":"Hello"}]}'
Sources: https://api-docs.deepseek.com/quick_start/pricing · https://api-docs.deepseek.com/updates · https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash