8–12 Sept 2026
Europe/Vienna timezone

Non-intrusive, deep-learning-based prediction of speech intelligibility and listening effort across diverse acoustic scenarios

FA2026/631
10 Sept 2026, 08:40
20m
Saal 10 (Messe Congress Graz)

Saal 10

Messe Congress Graz

A14 Physiological Acoustics and Audiology A14.08 Computational and AI approaches in audiology

Speaker

Hartmut Schoon (Carl von Ossietzky Universität Oldenburg)

Description

Models that predict auditory perception are important tools to quantify the effect of sound processing, for instance, speech enhancement algorithms in communication systems or in hearing aids. Non-intrusive models which do not rely on a clean reference signal could potentially be applied in real-world settings if they generalize well in broadly varying acoustic conditions. This contribution compares two non-intrusive deep-learning-based systems for predicting two metrics, subjective listening effort (LE) as well as speech intelligibility (SI). SI and LE predictions are compared in complex binaural scenes, based on speech enhanced by hearing aid algorithms, and synthetic speech. The PHOne-based Binaural Intelligibility model (PHOBI) predicts both metrics by measuring phone prediction uncertainties of an automatic speech recognition system. HASANet+ is a deep-learning framework trained to mimic the output of two intrusive speech perception models without a clean reference signal, utilizing the WavLM foundation model for acoustic feature extraction. Both models show strong correlations (r > 0.88) across all tasks, confirming the validity of non-intrusive approaches for speech perception modeling.

Authors

Hartmut Schoon (Carl von Ossietzky Universität Oldenburg) Dirk Eike Hoffner (Carl von Ossietzky Universität Oldenburg) Rainer Huber (Carl von Ossietzky Universität Oldenburg) Jan Rennies (Fraunhofer IDMT-HSA, Oldenburg) Bernd T. Meyer (Carl von Ossietzky Universität Oldenburg)

Presentation materials

There are no materials yet.