---
title: "NVIDIA releases open speech-to-speech AI model with 450ms latency"
url: https://www.parallelquant.com/posts/nvidia-releases-open-speech-to-speech-ai-model-with-450ms-latency-a55b48
source_name: "MarkTechPost"
source_url: https://www.marktechpost.com/2026/08/09/nvidia-releases-nemotronlabs-voicechat-11b-an-open-full-duplex-speech-to-speech-model-with-450-ms-turn-taking-and-live-tool-calling/
published: 2026-08-09T23:58:34.000Z
topics: ["open source"]
publisher: "Parallel Quant"
---

# NVIDIA releases open speech-to-speech AI model with 450ms latency

*2026-08-09 · Source: [MarkTechPost](https://www.marktechpost.com/2026/08/09/nvidia-releases-nemotronlabs-voicechat-11b-an-open-full-duplex-speech-to-speech-model-with-450-ms-turn-taking-and-live-tool-calling/)*

NVIDIA released NemotronLabs VoiceChat 11B, an open-weight full-duplex speech-to-speech model with about 450 milliseconds of turn-taking latency. The model supports live tool calling during voice conversations, letting it invoke external functions mid-dialogue.

**Why it matters:** Sub-500ms latency with live tool calling closes much of the gap between open models and proprietary voice assistants, making real-time voice agents more feasible to self-host. It adds to a wave of open-weight releases pressuring closed labs on cost and licensing terms for voice and agentic use cases.

**Topics:** open source

---
Read the original: https://www.marktechpost.com/2026/08/09/nvidia-releases-nemotronlabs-voicechat-11b-an-open-full-duplex-speech-to-speech-model-with-450-ms-turn-taking-and-live-tool-calling/
Canonical: https://www.parallelquant.com/posts/nvidia-releases-open-speech-to-speech-ai-model-with-450ms-latency-a55b48
Published by Parallel Quant — https://www.parallelquant.com
