8–12 Sept 2026
Europe/Vienna timezone

Multi-Teacher Distillation for Domain-Robust, Real-Time Representations of Electro-laryngeal Speech

FA2026/648
8 Sept 2026, 14:00
20m
Saal 5 (Messe Congress Graz)

Saal 5

Messe Congress Graz

Speaker

Benedikt Mayrhofer (Signal Processing and Speech Communication Laboratory)

Description

Recent advances in self-supervised learning (SSL) have significantly improved speech representations, yet their performance degrades in pathological domains such as electrolaryngeal (EL) speech. Additionally, the large computational footprint of state-of-the-art SSL models limits their applicability in real-time, on-device voice rehabilitation systems. We propose a multi-teacher knowledge distillation framework to train a lightweight, fully causal, and streaming-compatible content encoder that generalizes across healthy (HE) and EL speech. Our approach leverages two complementary teachers: (1) a frozen SSL model that provides phonetic cluster targets derived from healthy speech, and (2) an EL-adapted ASR model that supplies bottleneck feature regression targets to anchor representations in the pathological domain. Experimental results, evaluated via downstream ASR using word error rate (WER) and character error rate (CER), show that the proposed method substantially reduces error rates on EL speech compared to zero-shot SSL baselines, while maintaining competitive performance on HE data. We further analyze the impact of temporal modeling by comparing causal CNN, Transformer, Conformer, and Mamba-based student architectures, showing that effective temporal context modeling is a key factor for cross-domain generalization. The final model operates in real time on a single CPU core, providing a compact and practical representation backbone for cross-domain voice conversion.

Authors

Benedikt Mayrhofer (Signal Processing and Speech Communication Laboratory) Franz Pernkopf (Signal Processing and Speech Communication Laboratory) Philipp Aichinger (Medical University of Vienna) Martin Hagmüller (Signal Processing and Speech Communication Laboratory)

Presentation materials

There are no materials yet.