Phi-4-Reasoning-Vision-15B
Newest Phi model found (Mar 2026). Foundry model id and pricing not verified. Microsoft MAI models (MAI-Image-2/2.5, MAI-Voice-2, MAI-Transcribe-2, MAI-Thinking-1) are in Foundry but not covered by a file here.
- Context window
- 16,384 tokens
- Input
- text, image
- Output
- text
- License
- mit
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Hugging Face | — | huggingface.co/microsoft/Phi-4-reasoning-vision-15B | — |
| Microsoft Foundry | — | aka.ms/Phi-4-r-v-foundry | docs |
Notable capabilities (2)
- Hybrid think / no-think vision reasoning: Automatically chooses direct answers for perception tasks and long chain-of-thought only for math/science/diagram problems. source
- GUI grounding for computer-use agents: Dynamic-resolution SigLIP-2 encoder (up to 3,600 visual tokens) with strengths in GUI grounding for computer-use agents. source
Compact open multimodal reasoning model for visual math/science and screen understanding.
from transformers import AutoProcessor, AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained("microsoft/Phi-4-reasoning-vision-15B", trust_remote_code=True, device_map="auto")
Sources: https://huggingface.co/microsoft/Phi-4-reasoning-vision-15B
Other Microsoft models
VibeVoice (ASR, ASR-Streaming, ASR-BitNet, Realtime-0.5B TTS) · MAI-Transcribe-2 · MAI-Voice-2 / MAI-Voice-2-Flash · Phi-4 (14B)