Speaker
Description
Passive acoustic monitoring (PAM) has become a widely used tool for biodiversity assessment, particularly for monitoring vocalising animal species. However, in soundscapes dominated by natural or urban background noise, or in the presence of human voices, reliable detection and classification of animal vocalisations becomes challenging for widely used pretrained classifiers, such as BirdNET, and even expert auditory analysis. Conversely, when assessing the impact of anthropogenic noise such as the noise of high-altitude aircraft, intense biological activity—such as bird dawn chorus—can interfere with the monitoring.To address these challenges, we investigate artificial intelligence–based sound separation as a preprocessing step for both biodiversity monitoring and environmental noise assessment. Two approaches are explored. The first approach, inspired by speech enhancement techniques, employs complex ideal ratio masks trained with a U Net architecture to separate bird vocalisations from natural or urban background sounds. Results show that this effectively enhances bird sound separation and improves the accuracy of automatic classification.The second approach is based on a recently proposed transformer based Task-aware Unified Source Separation (TUSS), which allows token controlled sound extraction. By fine tuning the pretrained model for task specific targets, such as aircraft noise or animal vocalisations, the system accuracy increased. When applied to aircraft noise, the separated sound can be directly used to estimate its contribution to overall noise levels. When applied to animal vocalisations, it can serve as an enhanced input for subsequent PAM.These results demonstrate that AI based sound separation can support a more robust assessment of the effect of anthropogenic noise on biodiversity in complex real world soundscapes.