Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions

2026-08-10Sound

SoundArtificial IntelligenceComputation and Language
AI summary

The authors studied how well automated systems can judge the quality of computer-generated speech (Text-to-Speech or TTS). They broke down the idea of "naturalness" in speech into 10 specific aspects and had experts rate 860 speech samples. They tested four common automated systems that predict human ratings and found these systems mainly focus on sound quality and miss many other important speech details. The authors provide their data and tools to help improve how machines evaluate speech more accurately in the future.

Text-to-Speech (TTS)Mean Opinion Score (MOS)Audio Large Language Models (Audio-LLM)speech naturalnessspeech evaluationacoustic signal qualitylinguistic annotationperceptual dimensionsmeta-evaluation benchmarkautomated speech rating
Authors
Oluwanifemi Bamgbose, Simon Rosen, Jash Shah, Lindsay Devon Brin, Hoang H Nguyen, Anke Koelzer, Rachel Hansen, Tara Bogavelli, Fanny Riols
Abstract
Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive. We deconstruct "naturalness" into a linguistically grounded annotation schema spanning 10 distinct perceptual dimensions, and use it to construct the first dimension-level meta-evaluation benchmark for TTS, comprising 860 utterances annotated by trained linguist raters. Results from benchmarking four MOS predictors and four Audio-LLM judges reveal that MOS predictors collapse onto acoustic signal quality, while Audio-LLM judges show selective, prompt-dependent detection that does not generalise across all dimensions. Neither class reliably captures a breadth of linguistically structured speech errors. Our dataset, annotation schema, and evaluation code are publicly released to support more targeted and interpretable TTS evaluation.