8โ€“12 Sept 2026
Europe/Vienna timezone

Session

A20.00 Speech

A20.00
9 Sept 2026, 08:40

Conveners

A20.00 Speech: S144

  • Philipp Aichinger (Medical University of Vienna)
  • Franz Pernkopf (Signal Processing and Speech Communication Laboratory)
  • Peter Balazs (Acoustics Research Institute, Austrian Academy of Sciences)
  • Barbara Schuppler (Signal Processing and Speech Communication Laboratory)
  • Oliver Niebuhr (University of Southern Denmark)

A20.00 Speech: S288

  • Barbara Schuppler (Signal Processing and Speech Communication Laboratory)
  • Philipp Aichinger (Medical University of Vienna)
  • Oliver Niebuhr (University of Southern Denmark)
  • Peter Balazs (Acoustics Research Institute, Austrian Academy of Sciences)
  • Franz Pernkopf (Signal Processing and Speech Communication Laboratory)

A20.00 Speech: P438

  • Barbara Schuppler (Signal Processing and Speech Communication Laboratory)
  • Oliver Niebuhr (University of Southern Denmark)
  • Philipp Aichinger (Medical University of Vienna)
  • Franz Pernkopf (Signal Processing and Speech Communication Laboratory)
  • Peter Balazs (Acoustics Research Institute, Austrian Academy of Sciences)

Presentation materials

There are no materials yet.

  1. Mauro Manucci (ENS Paris)
    09/09/2026, 08:40
    A20 Speech

    Vowel recognition relies on specific physical properties of sound, known as ``acoustic cues'', that listeners use to identify phonemes. The first two formant frequency values (F1 and F2) have been repeatedly shown to be the primary acoustic cues for vowel categorisation. However, identifying the exact distribution underlying their mental representation remains a difficult task. Current...

    Go to contribution page
  2. Steve Gรถring (Audiovisual Technology Group, TU Ilmenau)
    09/09/2026, 09:00
    A20 Speech

    Text-to-speech systems are an essential building block in our modern AI-driven world. Various models have been established and are often also provided as open source. In this study, we evaluate ten different AI-based text-to-speech systems, with eight of them se- lected from the publicly available toolbox Coqui.ai TTS. In addition, with OpenAI and Google translate, two proprietary models are...

    Go to contribution page
  3. Azal LE BAGOUSSE (Laboratoire des Systรจmes Perceptifs, ENS-PSL)
    09/09/2026, 09:20
    A20 Speech

    Speech comprehension relies on normalization mechanisms that help listeners cope with the natural variability in speech production. In speech rate normalization, for instance, the perception of an ambiguous target phoneme can be influenced by the rate of the preceding context. While previous research has demonstrated that both proximal and distal contexts contribute to the effect, they provide...

    Go to contribution page
  4. Robert Gogol (Adam Mickiewicz University in Poznan)
    09/09/2026, 09:40
    A20 Speech

    This study explores the 'Speech-to-Song Illusion' discovered by Diana Deutsch. Specifically, it questions whether a correlation can be established between subjective human perception of the illusion and objective analyses performed by machine learning tools within the open-source Essentia library and others. To identify and examine the specific region, where machine learning algorithms yield...

    Go to contribution page
  5. Kivanc Kitapci (TOBB ETU)
    09/09/2026, 10:20
    A20 Speech

    This study investigated the effects of Spatial Release from Masking (SRM) on speech intelligibility in normal-hearing native Turkish speakers using the Turkish Modified Rhyme Test (MRT). Thirty participants (15 females, 15 males; mean age: 23.0 years, SD = 2.3) completed the MRT under nine spatial conditions defined by the combination of target word azimuth angle (0ยฐ, -50ยฐ, +50ยฐ) and...

    Go to contribution page
  6. Paolo Mesiano (Eriksholm Research Centre)
    09/09/2026, 10:40
    A20 Speech

    When conducting conversations in adverse acoustic conditions (e.g., in background noise), talkers tend to alter their speech production to increase speech clarity for their interlocutors. Compared to casual speech, such โ€œclear speechโ€ has been reported to facilitate speech intelligibility in challenging listening situations, but this evidence has been obtained through passive listening tests...

    Go to contribution page
  7. Daniel Aalto (University of Alberta)
    09/09/2026, 11:00
    A20 Speech

    Malocclusion impacts articulation of speech sounds, in particular fricatives. The present study compares spectral and prosodic patterns between adults with and without jaw disharmony during production of the English sibilants /s/ and /สƒ/. 36 adult participants were recruited and categorized using the Angle Classification into control (Class I) and experimental (Class II and Class III) groups....

    Go to contribution page
  8. Lionel Feugรจre (Bureau d'Enquรชtes et d'Analyses (BEA))
    09/09/2026, 11:20
    A20 Speech

    Aircraft Cockpit Voice Recorders (CVRs) contain sensitive data which are incompatible with open database to develop specific speech processing tools. The access to aircraft is also difficult and expensive. For these reasons, no realistic CVR database exists to be shared with the scientific community to train their algorithms. Acoustic and electronic measurement campaigns were conducted in...

    Go to contribution page
  9. Etienne Gaudrain (CRNL, CNRS UMR5292, Inserm U1028, Universitรฉ Lyon 1)
    09/09/2026, 11:40
    A20 Speech

    Previous studies have demonstrated that vocal tract length (VTL) perception is strongly limited in most cochlear-implant users, with post-lingually deaf adults showing greater deficits than those implanted in early childhood. Three potential factors may underlie this difficulty, as well as this difference between groups: (1) the implantโ€™s coding strategy, which may distort VTL cues as it is...

    Go to contribution page
  10. Barbara Schuppler (Signal Processing and Speech Communication Laboratory)
    09/09/2026, 15:00
    A20 Speech

    Conversational overlaps, i.e., instances where two or more participants speak simultaneously, are a prevalent feature of human communication, influencing conversational dynamics and speaker roles. While previous studies have mainly analysed and categorised them based on competitiveness, this paper focuses on the roles of individual speakers within these overlaps. Our study is based on...

    Go to contribution page
  11. Hansjรถrg Mixdorff (Berliner Hochschule fรผr Technik)
    09/09/2026, 15:00
    A20 Speech

    Isaฤenko and Schรคdlich describe German F0 contours as communicatively motivated tone switches aligned with accented syllables. Stock and Zacharias link these switches to phonologically distinct intonational elements, or intonemes, with three main classes distinguished by sentence modality: information, contact and non-terminal intoneme. Using the Fujisaki model, tone switch timing is captured...

    Go to contribution page
  12. Barbara Schuppler (Signal Processing and Speech Communication Laboratory)
    09/09/2026, 15:00
    A20 Speech

    This study investigates the perception of disfluencies in speech, focusing on the fillers aehm" andalso". Using conversational speech recordings from the Austrian German GRASS corpus, we select 72 stimuli similiarly long in number of tokens that contained either aehm" oralso" and conduct a perception experiment with 30 participants. During the experiment, listeners rated for each...

    Go to contribution page
  13. Hansjรถrg Mixdorff (Berliner Hochschule fรผr Technik)
    09/09/2026, 15:00
    A20 Speech

    This analysis of the link between speech rate and speech reduction identifies the degree to which rate alone drives reduction. Using the German part of the Bonn-Tempo corpus, we analyze vowel, consonant, and pause reduction as a function of speech rate and prosodic accentuation. Phone- and syllable-based rate measures were first computed, followed by Pfitzingerโ€™s perceptual local speech...

    Go to contribution page
Building timetable...