VibeVoice (ASR, ASR-Streaming, ASR-BitNet, Realtime-0.5B TTS)
Open-source voice research family from Microsoft (MIT). Timeline: TTS 2025-08-25 (code pulled 2025-09-05), Realtime-0.5B streaming TTS (~300 ms first audio) 2025-12-03, ASR 2026-01-21, Transformers integration 2026-03, Foundry Labs 2026-03-12, ASR-BitNet 2026-07-23, ASR-Streaming (10 languages, hotwords, speaker attribution) announced 2026-09-03; HF repos microsoft/VibeVoice-ASR-Streaming-7B and -1.5B created 2026-09-02 (verified 2026-09-29). Monthly downloads to 2026-09-29: VibeVoice-ASR ~734k, VibeVoice-1.5B ~717k. Separate from Microsoft's proprietary MAI-Voice/MAI-Transcribe.
- Input
- audio, text
- Output
- text, audio
- License
- mit
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Hugging Face | microsoft/VibeVoice-ASR | huggingface.co/microsoft/VibeVoice-ASR | — |
| Hugging Face (Transformers format) | microsoft/VibeVoice-ASR-HF | huggingface.co/microsoft/VibeVoice-ASR-HF | — |
| Hugging Face | microsoft/VibeVoice-ASR-BitNet | huggingface.co/microsoft/VibeVoice-ASR-BitNet | — |
| Hugging Face | microsoft/VibeVoice-Realtime-0.5B | huggingface.co/microsoft/VibeVoice-Realtime-0.5B | — |
| Hugging Face (streaming ASR, 7B repo; 9B params incl. decoder) | microsoft/VibeVoice-ASR-Streaming-7B | huggingface.co/microsoft/VibeVoice-ASR-Streaming-7B | — |
| Hugging Face (streaming ASR, small) | microsoft/VibeVoice-ASR-Streaming-1.5B | huggingface.co/microsoft/VibeVoice-ASR-Streaming-1.5B | — |
| GitHub | — | github.com/microsoft/VibeVoice | — |
Notable capabilities (4)
- 60-minute single-pass ASR with diarization: VibeVoice-ASR (~9B params incl. Qwen2-based decoder) transcribes up to 60 min in one pass with who/when/what structured output, hotwords and 50+ languages with code-switching. source
- CPU-only realtime ASR (found after launch): VibeVoice-ASR-BitNet (2026-07-23) compresses the model 4.62 GB -> 1.58 GB and runs faster than real time on 3 CPU threads (1.6-2.3x faster than Whisper.cpp). source
- Streaming speaker-attributed ASR (found after launch): VibeVoice-ASR-Streaming (7B and 1.5B repos, uploaded 2026-09-02) transcribes live audio with speaker attribution (who said what) and custom hotwords in 10 languages (zh, en, fr, de, it, ja, ko, pt, ru, es); MIT license. source
- Long-form multi-speaker TTS (withdrawn) (found after launch): Original VibeVoice-TTS (1.5B/7B) generated up to 90 min with 4 speakers; Microsoft removed the TTS code on 2025-09-05 over responsible-AI misuse concerns. source
Open Microsoft speech models for long-form ASR and lightweight streaming TTS.
See the Hugging Face Transformers docs (model_doc/vibevoice_asr) and the GitHub repo for inference code.
Sources: https://github.com/microsoft/VibeVoice · https://huggingface.co/microsoft/VibeVoice-ASR · https://huggingface.co/docs/transformers/model_doc/vibevoice_asr
Other Microsoft models
Phi-4-Reasoning-Vision-15B · MAI-Transcribe-2 · MAI-Voice-2 / MAI-Voice-2-Flash · Phi-4 (14B)