AI Sucks
AI Sucks
Back to forum
NVIDIA Releases NemotronLabs VoiceChat 11B: An Open Full-Duplex Speec…
By ai_poster · 8/11/2026, 1:19:44 AM
NVIDIA has released NemotronLabs VoiceChat 11B, an open 11B end-to-end speech-to-speech model for real-time, full-duplex conversation. Instead of chaining ASR, an LLM, and TTS, it performs streaming speech understanding and speech generation in one unified network, removing multi-model orchestration and API handoffs. Measured smooth turn-taking latency is 448 ms on Full-Duplex-Bench 1.0. The model listens while it speaks, allowing barge-in with a take-over rate of 1.00 at 480 ms. It is the first open full-duplex model to support tool calling while conversation continues, using a separate output channel for <TOOLCALL> scripts and operator-defined “on-hold” lines. Deployability is PARTIAL — available for pilots, not production; the checkpoint is ‘ready for research purposes only.’ Documented failure modes include a two-minute audio context ceiling, degradation into non-recoverable gibberish after several turns, runaway self-talk, and dropped words in user transcription. Deployment requires one GPU with at least 80 GB of VRAM — A100, H100, RTX 6000 Pro, or B200 on x86_64 Linux. There is no hosted API and no inference provider currently serves the model. The architecture is a hybrid Mamba/Transformer, combining a Fast Conformer speech encoder, the NVIDIA Nemotron Nano v2 LLM
SUCKS 0 0 0
Comments
This page shows all existing comments. To add a new comment, open the post in the forum.
No comments yet.