Post-Cutoff.com
  1. Home
  2. Timeline
  3. 2026
  4. NVIDIA releases NemotronLabs VoiceChat 11B, an open…

NVIDIA releases NemotronLabs VoiceChat 11B, an open full-duplex voice model with tool calling

★★★after cutoffopen-sourceNVIDIAconfidence: medium

NVIDIA published NVIDIA-NemotronLabs-VoiceChat-11B on Hugging Face (card dated 2026-08-03; arXiv 2609.21967): an end-to-end, full-duplex speech-to-speech model (FastConformer encoder + Nemotron Nano v2 9B + TTS decoder) that NVIDIA calls the first open full-duplex model to support tool calling. It has ~450 ms turn-taking latency and ranks #2 among open models on VoiceBench and Full-Duplex-Bench.

Key facts

What happened

NVIDIA added an 11B end-to-end full-duplex voice model to its Nemotron Speech collection. It listens and speaks at the same time, and it can call external tools while keeping the conversation going. Before this, open full-duplex models did not do tool calling.

Why it matters

Open full-duplex models (Kyutai Moshi, NVIDIA PersonaPlex) were mostly chat demos. Tool calling makes an open, self-hostable alternative to cascaded ASR→LLM→TTS agents and to closed realtime APIs possible. The "first" is NVIDIA's own claim. The release date comes from the model card; we found no separate press release.

Changelog

  • 2026-09-29: created

Models

Related events

  1. Google ships Gemini 3.8 Live voice models and Gemini 3.8 Flash TTS with voice design and cloning ★★
  2. OpenAI launches GPT-Live, full-duplex voice models replacing ChatGPT's Advanced Voice Mode ★★★★

Sources (3)

id: 2026-08-03-nvidia-nemotronlabs-voicechat · updated 2026-09-29 · open in the interactive timeline