Muse Voice Transcribe 1.0
Meta's first real-time audio perception model on the Meta Model API (launched 2026-09-03); 25+ languages. Speech-to-text only: Meta does not offer a TTS or speech-to-speech API; Muse's realtime voice mode and Muse Realtime Avatar (Connect, 2026-09-23) are consumer features without a documented API.
- Input
- audio
- Output
- text
- License
- proprietary
- Pricing
- per 1k minutes: $3 (USD per 1,000 minutes of audio ($0.18/hour)) source
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Meta Model API (streaming) | muse-voice-transcribe-1.0 | wss://api.meta.ai/v1/asr/realtime | docs |
| Meta Model API (file) | muse-voice-transcribe-1.0 | https://api.meta.ai/v1/asr/transcribe | docs |
Notable capabilities (2)
- #1 streaming STT on Artificial Analysis (claimed): Meta says it ranks first on the Artificial Analysis streaming speech-to-text leaderboard and had the lowest average diarization error rate among APIs tested, streaming and offline. source
- Diarization, VAD and endpointing in one model: Speaker attribution for 20+ speakers, punctuation, speech-boundary detection and adaptive delay (uses more audio context only for ambiguous words). source
Streaming and file transcription for voice agents built on Muse.
Sources: https://dev.meta.ai/resources/blog/meet-muse-voice-transcribe-streaming-speech-to-text/ · https://dev.meta.ai/docs/overview
Timeline entry
- Meta launches Muse Voice Transcribe, its first real-time speech model on the Meta Model API ★★★
On 2026-09-03 Meta Superintelligence Labs released Muse Voice Transcribe (muse-voice-transcribe-1.0), a streaming and file speech-to-text model on the Meta Model API at $0.18/hour that Meta says ranks #1 on the Artificial Analysis streaming STT leaderboard, with built-in diarization for 20+…
Other Meta models
Muse Spark 1.3 · Muse Glimmer 30B · Omnilingual ASR · Llama 4 Maverick (17B-128E) · Llama 4 Scout (17B-16E)