8–12 Sept 2026
Europe/Vienna timezone

Text-to-Speech Data Augmentation for Dysarthric Child Speech Reconstruction

FA2026/227
8 Sept 2026, 13:40
20m
Saal 5 (Messe Congress Graz)

Saal 5

Messe Congress Graz

Speaker

Moritz Pfeiler (Signal Processing and Speech Communication Laboratory)

Description

The goal of Dysarthric Speech Reconstruction is to improve intelligibility by converting dysarthric into healthy speech. In the future, this technology could become a valuable communication aid for people with this speech impairment. The development of systems performing complex speech tasks typically requires large amounts of data, which is especially sparse for pathological child speech. This paper explores the potential of Text-to-Speech data augmentation to improve performance in resource constraint scenarios. In this augmentation technique, synthetic dysarthric child speech is generated from text. To evaluate the effectiveness, a Voice Conversion model is trained with just 2.37 hours of real speech in two configurations: (1) only real speech (2) pre-trained with synthetic speech and fine-tuned on real speech. While the overall intelligibility remains low, data augmentation leads to noticeable improvements, which encourages further research. Future work should focus on optimizing both Text-to-Speech and Voice Conversion models for this task.

Authors

Moritz Pfeiler (Signal Processing and Speech Communication Laboratory) Benedikt Mayrhofer (Signal Processing and Speech Communication Laboratory) Barbara Schuppler (Signal Processing and Speech Communication Laboratory) Martin Hagmüller (Signal Processing and Speech Communication Laboratory)

Presentation materials