Speaker
Description
Hearing research has studied the acoustic basis of communication extensively. However, speech understanding is not purely an acoustic process: visual information from the speaker’s face, especially lip movements, plays a key complementary role. These cues become particularly valuable in challenging or noisy listening conditions. As virtual environments expand in both research and everyday communication, digital avatars are increasingly used not only as a medium for interaction but also as a tool for investigating human communication processes. Recent studies suggest that speech intelligibility in such contexts strongly depends on the animation methods used to drive avatars, yet remains lower than with natural video recordings. One major reason is that the visual speech cues produced by avatars often lack the precision required to reliably discriminate certain phonetic elements.This work presents a pipeline for the animation and control of realistic avatars in virtual environments. The proposed approach relies on video recordings as reference material for the production of facial animations, rather than on direct or fully automated methods, allowing for a more precise treatment of complex visemes. The pipeline is designed in a modular way to integrate within our virtual environment. Within this context, particular attention is given to issues of audio-visual alignment, especially the temporal relationship between speech and lip movements. Finally, we offer initial insights into the potential benefits of such animations for speech intelligibility.