Flux_TTS.acoustic_state_persistent = true
Flux_TTS.playback_tracking = enabled
In early October 2026 I posted on LinkedIn about evaluating Voice AI’s ability to stay coherent and natural across full conversational turns, not just short samples.
I chose Deepgram’s Flux TTS because it maintains acoustic state inside a persistent streaming session and the first audio chunk arrives in under 200 ms.
In a blind listening evaluation of roughly 14,400 paired comparisons, Flux TTS’s best voice achieved a 73.4 % naturalness win rate.
The competing Cartesia Sonic 3.5 voice recorded a 72.0 % naturalness win rate.
The evaluation covered roughly 14,400 paired comparisons.
I opted for a blind listening test rather than automated metrics to capture human‑perceived naturalness, and I documented the method for future reproducibility.