Cartesia has released Sonic-3.6, a streaming text-to-speech model that now ranks #1 on both Artificial Analysis speech leaderboards — 1,283 Elo on Provider Voice and 1,123 on Controlled Voice, the board that clones every model onto the same eight reference voices to isolate the synthesis engine. The model is built on state space models rather than transformers, and Cartesia states sub-90ms time-to-first-audio. It is available in beta on Cartesia's own API.

For decision-makers, this is more than a benchmark bump. The Controlled Voice result is the one that matters: by forcing every model onto identical reference voices, it strips away the advantage of a carefully chosen speaker and measures the raw quality of the synthesis engine. Leading there means Sonic-3.6 delivers consistent, high-quality speech across voices — not just one polished demo. Combined with sub-90ms latency, this pushes the practical boundary for real-time voice applications: assistants, customer service, and interactive agents where responsiveness is a feature, not a footnote.

The architectural choice is also strategic. State space models are a different path from the transformer-based approach most TTS systems use, and Cartesia's lead suggests that this alternative can compete — or win — on quality and speed. For teams building or buying voice AI, that widens the option space and puts pressure on incumbents to match the latency and quality bar. The beta availability on Cartesia's API means it's not just a research claim; it's a product you can evaluate today.