Speaker
Description
Numerically simulated Head-Related Transfer Functions (HRTFs) provide an attractive alternative to acoustic measurement, removing the need for anechoic facilities and dedicated measurement equipment. However, their perceptual fidelity remains limited compared to gold-standard acoustic measurements, with residual errors in interaural time differences (ITDs), interaural level differences (ILDs) and monaural spectral, cues which govern lateral and elevation localisation. This seems to be mainly (but not only) due to the underlying mesh acquisition method or the simulation pipeline, which result in lower performances when assessed through perceptual models. The aim of this study is to provide a machine learning based post-processing tool that refines synthetic HRTFs towards acoustically measured fidelity, independently of the mesh source used. To this end, neural network architectures are trained on the Extended SONICOM HRTF dataset (200 subjects), each providing paired acoustically measured HRTFs alongside two synthetic counterparts obtained from a high-resolution three-dimensional scan and from a photogrammetry-reconstructed (PR) mesh, enabling refinement to be learnt across heterogeneous synthesis inputs. Networks are trained subject-independently using different perceptual and non-perceptual loss functions, and evaluated on unseen subjects through numerical metrics (ITD, ILD, log-spectral distortion), auditory-model predictions and a behavioural VR sound localisation test. Improvements are found in each evaluation method, supporting neural refinement as a simple post-processing stage that compensates for synthesis artefacts and brings simulated HRTFs closer to measurement-grade fidelity.