parallelquant
August 18, 2026 · MarkTechPost

Cartesia's new TTS model tops both speech leaderboards

Cartesia released Sonic-3.6, a streaming text-to-speech model built on state space models instead of transformers. It now ranks #1 on both Artificial Analysis speech arenas, with sub-90 millisecond time-to-first-audio, and is available in beta on Cartesia's API.

Why it matters: The result adds to evidence that state-space architectures can outperform transformers for latency-sensitive tasks like real-time voice, a niche where response speed matters more than raw scale. Sub-90ms first-audio latency pushes streaming TTS closer to feeling truly conversational, relevant for voice agents and real-time assistants.

Related updates