Qwen3.8-27B
Best Apache-2.0 Qwen for self-hosting; also the go-to open Qwen VL model (Qwen3-VL successor). First-party hosted API 'coming soon' on Qwen Cloud at time of check. Pricing not verified (no first-party price).
- Context window
- 262,144 tokens
- Input
- text, image, video
- Output
- text
- License
- apache-2.0
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Hugging Face | — | huggingface.co/Qwen/Qwen3.8-27B | — |
| Hugging Face (FP8) | — | huggingface.co/Qwen/Qwen3.8-27B-FP8 | — |
| OpenRouter | qwen/qwen3.8-27b | openrouter.ai/qwen/qwen3.8-27b | — |
| OpenRouter (free tier) | qwen/qwen3.8-27b:free | openrouter.ai/qwen/qwen3.8-27b | — |
| Web app | — | chat.qwen.ai | — |
Notable capabilities (3)
- Dense open VLM with agentic focus: 27B dense native vision-language model (images and hour-scale video) tuned for coding and long-horizon agent tasks, Apache-2.0. source
- Thinking control: Thinking on by default, can be disabled per request; reasoning_effort and preserve_thinking supported. source
- Extensible to 1M context: 262,144 tokens native, extensible up to 1,000,000. source
Open-weight (Apache-2.0) dense multimodal model for local/self-hosted coding agents and vision tasks; runs on vLLM, SGLang, Transformers.
vllm serve Qwen/Qwen3.8-27B --max-model-len 262144
Sources: https://huggingface.co/Qwen/Qwen3.8-27B · https://openrouter.ai/qwen/qwen3.8-27b
Other Alibaba (Qwen) models
Qwen-Audio-3.1-ASR (Flash) · Qwen-Audio-3.1-Realtime (Plus) · Qwen-Audio-3.1-TTS-Next · Qwen3.8-LiveTranslate (Flash Realtime) · Qwen3.8-Omni-Flash · Qwen3.8-Flash · Qwen3.8-Max · Qwen-Image-3.0 (Pro) · Qwen3.7-Plus · Qwen3-ASR (0.6B / 1.7B) + Qwen3-ForcedAligner · Qwen3-TTS (open weights 0.6B / 1.7B; API qwen3-tts-flash)