Speaker
Description
Perception of a soundscape is shaped significantly by the dominant sound sources present in an acoustic environment. Spatial characteristics have also been suggested to influence how a soundscape is perceived. Soundscape recordings are often analysed holistically rather than at the level of individual sources; the ability to separate and analyse sources individually could enable deeper insight into how specific sources drive overall perception.Spatial audio formats are commonly used in soundscape assessment, yet a source stripped of its directional character cannot be used to investigate spatial contributions to perception, or re-rendered for listening experiments. The preservation of spatial characteristics, however, appears to be relatively unexplored for the task of universal sound separation.This paper presents a minimal modification of Conv-TasNet to accept and output first-order ambisonic (FOA) signals for universal sound separation (separating sources of any class), comparing three loss functions (SNR, SI-SDR, and a multichannel SI-SDR variant) for FOA-to-FOA separation. Results indicate that the multichannel SI-SDR loss substantially improves both separation performance and preservation of source direction compared with conventional SI-SDR and SNR losses. These early-stage findings suggest spatially-preserving source separation may hold promise as a tool for future soundscape studies.