Speaker
Description
Low-frequency sound reproduction in concert-halls can be significantly improved through the use of spatial sound field control methods. Such techniques may require accurate spatial sampling of the room transfer functions (RTFs) across the venue for every subwoofer loudspeaker. Respecting the Nyquist criteria spatially is impractical in real-world concert scenarios where measurement time is limited. To that matter, RTF reconstruction methods estimate RTFs from a reduced number of measurements. Data-driven approaches, including deep learning models trained on simulated datasets, have recently shown promises in this area. We study Flow Matching Neural Processes (FMNPs) for the RTF reconstruction task. FMNP integrates the flow matching paradigm into the neural process framework: a U-Net transformer encoder aggregates measurements at arbitrary spatial positions and generates RTFs by learning a velocity field transporting samples from an initial Gaussian distribution toward the RTF distribution. To train and compare FMNP alongside existing methods from the literature, we introduce a dataset of room impulse responses simulated with a GPU-accelerated finite difference time domain method on large concert hall geometries, with variability introduced through room deformations and varied absorption conditions. This dataset is used to evaluate the practical feasibility of data-driven RTF reconstruction at scale.