8–12 Sept 2026
Europe/Vienna timezone

Evaluating unsupervised clustering of deep embeddings for unknown species identification in bioacoustics

FA2026/668
8 Sept 2026, 16:00
20m
Saal 3 (Messe Congress Graz)

Saal 3

Messe Congress Graz

Speaker

Vincent S. Kather (Naturalis Biodiversity Center)

Description

Natural soundscapes, especially in hyper-diverse environments often contain unknown sound events. Strategies to identify unknown sound events, especially those produced by species not included in training sets have traditionally been infeasible to manually identify. Recently, the classification performance of bioacoustic deep learning models is becoming competitive with human annotators. Yet, classifier predictions are often unreliable for hyper-diverse soundscapes and under-represented regions or taxa, and for unknown species they are inherently ill-defined. Recent research suggests that by using the feature extraction capability of state-of-the-art bioacoustic deep learning models, similar sounds can be clustered together in the feature space despite the model never having encountered these sounds during training. However, current methods to evaluate the ability of bioacoustic feature extractors to specifically identify unknown sounds are limited. Therefore, in this study we curate a dataset of under-represented species vocalizations from diverse taxonomic groups (i.e. birds, mammals, amphibians and insects) and superimpose background noise from different habitats (i.e. tropical and temperate forests). By doing so we are able to vary signal-to-noise ratio (SNR) and background noise while mimicking real recorded soundscapes. We evaluate a wide variety of acoustic deep learning models and clustering algorithms. Preliminary results show that models trained on a large variety of species are capable of clustering sounds from species they have not been shown during training over different noise environments for high SNR values. This evaluation method provides a fine-grained understanding of the limitations of current deep learning models and, when paired with interactive visualizations, can be used to inspect how models handle unknown sounds.

Authors

Vincent S. Kather (Naturalis Biodiversity Center) Ben McEwen (Tilburg University) Sylvain Haupert (Museum Nacionale d’Histoire Naturell) Burooj Ghani (Naturalis Biodiversity Center) Dan Stowell (Naturalis Biodiversity Center)

Presentation materials

There are no materials yet.