8–12 Sept 2026
Europe/Vienna timezone

Evaluation of AI-based Text-to-Speech Generators

FA2026/165
9 Sept 2026, 09:00
20m
Saal 5 (Messe Congress Graz)

Saal 5

Messe Congress Graz

A20 Speech A20.00 Speech

Speaker

Steve Göring (Audiovisual Technology Group, TU Ilmenau)

Description

Text-to-speech systems are an essential building block in our modern AI-driven world. Various models have been established and are often also provided as open source. In this study, we evaluate ten different AI-based text-to-speech systems, with eight of them se- lected from the publicly available toolbox Coqui.ai TTS. In addition, with OpenAI and Google translate, two proprietary models are included for a wider evaluation. We conducted a perceptual evaluation based on ten pre-defined text-prompts. In an online study, participants rated audio quality, clarity, and naturalness. Our results indicate that the generators VITS and Google translate are the best regarding audio quality and clarity, while ChatGPT achieves the highest score in naturalness. In addition, a high correlation between quality, clarity, and naturalness could be observed. We further carried out an objective evaluation with low- and high-level audio features, where the Meta Audiobox Aesthetics shows the high- est correlation with all three dimensions. In the future, more generators and text-prompts should be included in similarly designed evaluation studies. Further, a comparison to real speech will be a valuable next step.

Authors

Steve Göring (Audiovisual Technology Group, TU Ilmenau) Usman Khalid (Audiovisual Technology Group, TU Ilmenau) Annika Neidhardt (Audio Engineering, Faculty of Media, HS Mittweida)

Presentation materials