8–12 Sept 2026
Europe/Vienna timezone

Multi-Modal Voice Conversion for Pathological Speech: Leveraging Visual Articulation to Restore a Healthy Voice

FA2026/938
8 Sept 2026, 14:00
3h
Messehalle (Poster+Exhibition) (Messe Congress Graz)

Messehalle (Poster+Exhibition)

Messe Congress Graz

Speaker

Enrique Orozco Olivares (Signal Processing and Speech Communication Laboratory)

Description

Pathological speakers whose intelligibility is caused by impairment of vocal fold movement retain largely intact lip and mouth movements, making the visual modality a reliable signal even when the acoustic channel is severely degraded. We propose a multimodal end-to-end voice conversion system that exploits this preserved articulatory information to convert pathological speech into a healthy-sounding waveform of the same utterance, without relying on a separate target speaker encoder or conversion module. Audio is encoded with Whisper and visual articulatory features with AV-HuBERT; a speech rate predictor dynamically adjusts the number of Q-Former query tokens, making the system robust to the reduced and variable speaking rates characteristic of pathological speakers. A Conformer processes the fused representation while injecting target speaker identity through cross-attention at every layer, and a HiFi-GAN vocoder synthesises the output waveform. Adversarial fine-tuning with a Constant-Q transform discriminator further improves perceptual quality. We hypothesize that the visual stream, capturing not only lip movements but also micro-expressions and head movements, improves the expressiveness of converted speech compared to audio-only approaches. Since no public parallel audio-visual dataset exists for voice conversion, training data is constructed synthetically at each stage using a combination of publicly available datasets, web-sourced material, and private pathological speech data recorded under controlled conditions. The current work targets pathological-to-healthy voice conversion in German, a task underrepresented in voice conversion research, nevertheless, the methodology is language-agnostic and transferable to other languages. Evaluation is across subjective naturalness and speaker similarity scores alongside objective speaker similarity metrics.

Authors

Enrique Orozco Olivares (Signal Processing and Speech Communication Laboratory) Benedikt Mayrhofer (Signal Processing and Speech Communication Laboratory) Philipp Aichinger (Medical University of Vienna) Martin Hagmüller (Signal Processing and Speech Communication Laboratory) Franz Pernkopf (Signal Processing and Speech Communication Laboratory)

Presentation materials

There are no materials yet.