Speaker
Description
Sound Source Characterization (SSC) with microphone arrays has seen significant advances through deep learning. However, most data-driven methods are trained on data from a single, fixed microphone array geometry, making the resulting models geometry-specific and difficult to transfer to other configurations. Moreover, many architectures do not explicitly use sensor position information. To overcome this limitation, a Transformer architecture that uses a Message Passing Neural Network for array-agnostic SSC is proposed. The proposed model operates on a graph constructed from spatial and spectral information of the microphones, from which the location and strength of the dominant sound source in the region of interest is estimated. The model is evaluated across multiple array geometry classes, such as spiral, grid and randomized layouts, and different numbers of channels outside the training distribution, demonstrating robust generalization where a fixed-geometry baseline fails entirely. Furthermore, training with a variable number of microphones is shown to improve performance. To the authors’ knowledge, this is the first work on grid-free SSC generalizing across large-scale array geometries of up to 128 microphones.