Speaker
Description
Pathological speakers whose intelligibility is caused by impairment of vocal fold movement retain largely intact lip and mouth movements, making the visual modality a reliable signal even when the acoustic channel is severely degraded. We propose a multimodal end-to-end voice conversion system that exploits this preserved articulatory information to convert pathological speech into a healthy-sounding waveform of the same utterance, without relying on a separate target speaker encoder or conversion module. Audio is encoded with Whisper and visual articulatory features with AV-HuBERT; a speech rate predictor dynamically adjusts the number of Q-Former query tokens, making the system robust to the reduced and variable speaking rates characteristic of pathological speakers. A Conformer processes the fused representation while injecting target speaker identity through cross-attention at every layer, and a HiFi-GAN vocoder synthesises the output waveform. Adversarial fine-tuning with a Constant-Q transform discriminator further improves perceptual quality. We hypothesize that the visual stream, capturing not only lip movements but also micro-expressions and head movements, improves the expressiveness of converted speech compared to audio-only approaches. Since no public parallel audio-visual dataset exists for voice conversion, training data is constructed synthetically at each stage using a combination of publicly available datasets, web-sourced material, and private pathological speech data recorded under controlled conditions. The current work targets pathological-to-healthy voice conversion in German, a task underrepresented in voice conversion research, nevertheless, the methodology is language-agnostic and transferable to other languages. Evaluation is across subjective naturalness and speaker similarity scores alongside objective speaker similarity metrics.