Speaker
Description
Natural soundscapes, especially in hyper-diverse environments often contain unknown sound events. Strategies to identify unknown sound events, especially those produced by species not included in training sets have traditionally been infeasible to manually identify. Recently, the classification performance of bioacoustic deep learning models is becoming competitive with human annotators. Yet, classifier predictions are often unreliable for hyper-diverse soundscapes and under-represented regions or taxa, and for unknown species they are inherently ill-defined. Recent research suggests that by using the feature extraction capability of state-of-the-art bioacoustic deep learning models, similar sounds can be clustered together in the feature space despite the model never having encountered these sounds during training. However, current methods to evaluate the ability of bioacoustic feature extractors to specifically identify unknown sounds are limited. Therefore, in this study we curate a dataset of under-represented species vocalizations from diverse taxonomic groups (i.e. birds, mammals, amphibians and insects) and superimpose background noise from different habitats (i.e. tropical and temperate forests). By doing so we are able to vary signal-to-noise ratio (SNR) and background noise while mimicking real recorded soundscapes. We evaluate a wide variety of acoustic deep learning models and clustering algorithms. Preliminary results show that models trained on a large variety of species are capable of clustering sounds from species they have not been shown during training over different noise environments for high SNR values. This evaluation method provides a fine-grained understanding of the limitations of current deep learning models and, when paired with interactive visualizations, can be used to inspect how models handle unknown sounds.