8–12 Sept 2026
Europe/Vienna timezone

Principles of auditory scene analysis emerge in computational audio separation

FA2026/844
9 Sept 2026, 15:00
20m
Halle A (Messe Congress Graz)

Halle A

Messe Congress Graz

Speaker

Bernhard U. Seeber (Technische Universität München)

Description

Auditory scene analysis describes the ability of the human auditory system to separate a mixture of sounds into the individual sources. The separation is based on bottom-up grouping information following Gestalt principles, such as harmonicity, onset synchrony and common modulation in amplitude and frequency, and learned top-down information. We investigated the separation mechanisms of a convolutional time-domain audio separation network (ConvTasNet). The network was trained for two-source separation on a variety of sounds: speech, environmental sounds, and music. The separation mechanisms were investigated by conducting a range of experiments on auditory scene analysis following classical experiments. The network learns to perform the separation and is able to separate the abstract experimental stimuli, thereby showing similar performance bounds as humans. The separation mechanisms exhibit the same principle organization rules as in the human auditory system: harmonicity, onset synchrony and common amplitude and frequency modulation. This suggests that machine learning of separation based on the statistics of sound stimuli results in similar mechanisms as have developed in humans.

Authors

Bernhard U. Seeber (Technische Universität München) Kean Chen (Northwestern Polytechnical University) Han Li (Technische Universität München)

Presentation materials

There are no materials yet.