Speaker
Description
The goal of Dysarthric Speech Reconstruction is to improve intelligibility by converting dysarthric into healthy speech. In the future, this technology could become a valuable communication aid for people with this speech impairment. The development of systems performing complex speech tasks typically requires large amounts of data, which is especially sparse for pathological child speech. This paper explores the potential of Text-to-Speech data augmentation to improve performance in resource constraint scenarios. In this augmentation technique, synthetic dysarthric child speech is generated from text. To evaluate the effectiveness, a Voice Conversion model is trained with just 2.37 hours of real speech in two configurations: (1) only real speech (2) pre-trained with synthetic speech and fine-tuned on real speech. While the overall intelligibility remains low, data augmentation leads to noticeable improvements, which encourages further research. Future work should focus on optimizing both Text-to-Speech and Voice Conversion models for this task.