August 9, 2026 · MarkTechPost
NVIDIA releases open speech-to-speech AI model with 450ms latency
NVIDIA released NemotronLabs VoiceChat 11B, an open-weight full-duplex speech-to-speech model with about 450 milliseconds of turn-taking latency. The model supports live tool calling during voice conversations, letting it invoke external functions mid-dialogue.
Why it matters: Sub-500ms latency with live tool calling closes much of the gap between open models and proprietary voice assistants, making real-time voice agents more feasible to self-host. It adds to a wave of open-weight releases pressuring closed labs on cost and licensing terms for voice and agentic use cases.