GLM-4.6V
Still sold on Z.ai (with FlashX and free Flash variants) but superseded by the natively multimodal GLM-5.3-Flash. Model id casing on Z.ai assumed lowercase glm-4.6v (listed as GLM-4.6V). Release day not verified (HF 2025-12-07).
- Context window
- 128,000 tokens
- Input
- text, image, video
- Output
- text
- License
- mit
- Pricing
- input: $0.3 · output: $0.9 · cache read: $0.05 (per 1M tokens (USD); GLM-4.6V-FlashX 0.04/0.4; GLM-4.6V-Flash free) source
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Z.ai API | glm-4.6v | https://api.z.ai/api/paas/v4 | docs |
| OpenRouter | z-ai/glm-4.6v | openrouter.ai/z-ai/glm-4.6v | — |
| Hugging Face | — | huggingface.co/zai-org/GLM-4.6V | — |
| Hugging Face (Flash) | — | huggingface.co/zai-org/GLM-4.6V-Flash | — |
| Web app | — | chat.z.ai | — |
Notable capabilities (2)
- Native multimodal function calling: First GLM vision model with native function calling (images can be passed to and returned from tools). source
- Interleaved image-text generation: Builds mixed image-text content from documents and tool-retrieved images; also frontend replication from screenshots. source
Open-weight (MIT) vision-language GLM; cheapest/free Z.ai vision option via GLM-4.6V-Flash.
Sources: https://huggingface.co/zai-org/GLM-4.6V · https://docs.z.ai/guides/overview/pricing