Speaker
Description
Two of the main challenges that the hearing care workforce currently face concern the limited resources and the increasing number of older adults who are in need of hearing screening. New technology tools, such as automatic speech recognition (ASR), have the potential to revolutionise the field. However, state-of-the-art ASR systems are predominantly trained on speech from healthy young adult native speakers. The ATHENA (Automating hearing assessments for all) project proposes the use of open-source ASR systems to automatically score four traditional speech audiometry tests: speech in quiet, speech in noise, speech in speech masker, and vocal emotion recognition.As a first step, four open-source ASR systems in Dutch (Kaldi NL, Whisper, Voxtral and NeMo) are evaluated when decoding participants’ spoken responses per speech audiometry test, focusing on the impact that decoding errors have on the test results. Three participant groups are included: native Dutch-speaking children, non-native Dutch-speaking adults, and native Dutch-speaking adults. As a second step, the ASR performance is tested with a wider group of non-native and native Dutch-speaking participants, including older adults of varying cognitive statuses, and hearing-impaired individuals. As a third step, we look into ASR customisation, including limiting the ASR lexicon per test (to words related to or included in the test speech material), alongside fine-tuning and data augmentation techniques. As a final step, the best-performing customised ASR systems coupled to the speech audiometry tests are implemented on engaging interfaces (such as socially assistive robots), and evaluated with respect to system robustness, reliability, and accessibility for clinical and non-clinical use.