Kyutai Moshi / Hibiki-Zero (full-duplex speech models)
Moshi (announced July 2024, weights + paper Sept 2024) is widely cited as the first real-time full-duplex open spoken dialogue model; NVIDIA PersonaPlex-7B (Jan 2026) is fine-tuned from Moshiko weights. Variants: moshiko (male)/moshika (female) in PyTorch bf16/int8, MLX int4/int8/bf16, Rust/Candle. Code MIT/Apache, weights CC-BY-4.0.
- Input
- audio
- Output
- audio, text
- License
- cc-by-4.0
- Verified
- 2026-09-29
How to call it
| Provider | Model id | Endpoint / URL | Docs |
|---|---|---|---|
| Hugging Face | kyutai/moshiko-pytorch-bf16 | github.com/kyutai-labs/moshi | — |
| Hugging Face | kyutai/hibiki-zero-3b-pytorch-bf16 | huggingface.co/kyutai/hibiki-zero-3b-pytorch-bf16 | — |
| Web demo | — | moshi.chat | — |
Notable capabilities (3)
- FIRST Open full-duplex spoken dialogue: 7B temporal transformer modelling user and Moshi audio streams simultaneously with an 'inner monologue' text stream; 160 ms theoretical / ~200 ms practical latency on an L4; Mimi codec (24 kHz, 12.5 Hz, 1.1 kbps). source
- Hibiki-Zero simultaneous speech translation (found after launch): 3B model (2026-02-12) translating French, Spanish, Portuguese and German speech to English in real time with voice transfer, trained without aligned data. source
- MoshiRAG (found after launch): Asynchronous knowledge retrieval via a text LLM for full-duplex speech models (2026-04-30); RL post-training for interactivity (2026-06-10). source
Sources: https://github.com/kyutai-labs/moshi , https://kyutai.org/blog/ , https://huggingface.co/kyutai/hibiki-zero-3b-pytorch-bf16