Speaker
Description
Text-to-speech systems are an essential building block in our modern AI-driven world. Various models have been established and are often also provided as open source. In this study, we evaluate ten different AI-based text-to-speech systems, with eight of them se- lected from the publicly available toolbox Coqui.ai TTS. In addition, with OpenAI and Google translate, two proprietary models are included for a wider evaluation. We conducted a perceptual evaluation based on ten pre-defined text-prompts. In an online study, participants rated audio quality, clarity, and naturalness. Our results indicate that the generators VITS and Google translate are the best regarding audio quality and clarity, while ChatGPT achieves the highest score in naturalness. In addition, a high correlation between quality, clarity, and naturalness could be observed. We further carried out an objective evaluation with low- and high-level audio features, where the Meta Audiobox Aesthetics shows the high- est correlation with all three dimensions. In the future, more generators and text-prompts should be included in similarly designed evaluation studies. Further, a comparison to real speech will be a valuable next step.