Not One Eval to Rule Them all: Contextualizing Synthetic Speech Perception Across Different Domains
Abstract
Text-to-speech (TTS) evaluation is an open challenge. While the primary target was "naturalness", recent fidelity gains shifted focus toward "appropriateness" and whether speech is correct for its context. In this work, we examine how perception changes when the expected downstream use varies. We measure the appropriateness and human-likeness of five SOTA TTS systems across five: AI assistant, reader, actor, animated character, and spontaneous speaker. Results show appropriateness varies across domains independent of naturalness. While systems excel at reading, expressive domains remain challenging, and optimizing for one can degrade others. Furthermore, naturalness scores tend to penalize stylized speech while rewarding spontaneity. Finally, our study also highlights blind spots in one-size-fits-all evaluation metrics across more expressive domains. We demonstrate that TTS expressivity is not "solved" but depends on the target domain, requiring context-aware evaluation.
Example Audio Samples from the Perception Study
| Speech Task | Transcript | GT | Elevenlabs v2 | Gemini TTS | Kokoro | GPT-4o-mini-tts | Kyutai TTS |
|---|---|---|---|---|---|---|---|
| Narration | Oh, if that horrid man had never come to Dalton! | ||||||
| Affect Conversational | Ew! What is that? Something exploded! | ||||||
| Spontaneous | I... I mean, they're just disgusting to me. | ||||||
| Inform | Would you like to review the progress you made on your goals today? |
Audio Labeling App