8–12 Sept 2026
Europe/Vienna timezone

Hydra VI-LoRA: Speaker-Aware Mixture-of-Experts for Personalized Recognition of Atypical Speech

FA2026/950
8 Sept 2026, 13:20
20m
Saal 5 (Messe Congress Graz)

Saal 5

Messe Congress Graz

Speaker

Niclas Pokel (ETH Zürich)

Description

Automatic speech recognition remains unreliable for individuals with impaired speech, in part because a single adapter must absorb the high inter-speaker acoustic variability that characterizes dysarthric, apraxic, and neurodegenerative populations. Parameter-efficient fine-tuning with Low-Rank Adaptation (LoRA), and its Bayesian extension via variational inference (VI-LoRA), can personalize foundation models such as Whisper to atypical speech, but a monolithic adapter still forces averaging across speakers whose articulation, prosody, and phonemic realization differ markedly. We extend VI-LoRA into a Hydra-style Mixture-of-Experts framework, Hydra VI-LoRA, in which multiple low-rank expert branches share a single down-projection but maintain independent stochastic up-projections, dispatched by a learned gating network. The router is trained without explicit speaker labels, allowing it to discover an implicit clustering of speakers and to adapt expert assignment dynamically at inference. We evaluate the framework on the English UA-Speech dysarthric corpus and the larger multi-etiology SAP-2024 dataset, with Whisper-Large V3 as the backbone. A per-layer top-k routing variant reduces character error rate on UA-Speech from 18% (pure LoRA) to 14%, and lowers WER on SAP to approximately 9% from a 10.5% LoRA baseline. Heatmap analyses reveal speaker- and etiology-correlated expert preferences that emerge without supervision, suggesting that the router captures meaningful subgroup structure. These preliminary results indicate that combining variational regularization with sparse expert routing is a promising path toward scalable, label-free personalization of ASR for atypical speech. We outline ongoing work on expert scaling, cross-lingual transfer, and interpretability of the discovered speaker clusters.

Author

Niclas Pokel (ETH Zürich)

Presentation materials