Speaker
Description
This paper focuses on the automatic speech recognition (ASR) of speech produced by elderly speakers diagnosed with dementia. In clinical contexts, reliable transcriptions can support the work of clinical linguists and serve as a basis for (automatic) analyses to identify linguistic biomarkers for dementia. ASR is further also increasingly used in voice-based assistive technologies supporting elderly people in everyday life. Since ASR systems are usually optimized for fluent, standard-language, they perform significantly worse on spontaneous speech from atypical speakers. This becomes even more challenging when dialectal variation and disfluencies are present. We investigate ASR methods for elderly Austrian German with a focus on faithful transcription rather than normalized text output. We compare four ASR architectures: zero-shot Whisper, fine-tuned Whisper model and two encoder-based models (wav2vec2 with greedy and with language-model decoding), which are first fine-tuned on adult Austrian German and then adapted to elderly speech. Besides standard evaluation metrics on word and character level, we present a detailed error analysis showing how well the models preserve disfluencies or hesitations and analyze transcription errors. Beyond Austrian German, this work provides methodological insights into ASR for atypical, disfluent speech in general, informing the development of more robust and clinically applicable speech technologies.