Speaker
Description
Accurate sound localisation in augmented reality depends also on employing head-related transfer functions (HRTFs) that match the listener's anatomy. Yet, predicting how a specific listener will perceive a given spatial audio rendering remains an open problem. Bayesian observer models can, in principle, predict individual localisation behaviour from acoustic measurements, but their parameters have traditionally been fitted by minimising ad hoc summary metrics, such as polar error and quadrant error rate. This metric-matching approach lacks a probabilistic foundation, prevents formal hypothesis testing, and reduces the full localisation response distribution to scalar summaries that discard the bimodal structure which distinguishes polar uncertainty from front-back confusions. We address this by reformulating parameter estimation as maximum-likelihood optimisation over full listener response distributions, yielding a model that is individually fitted and capable of tracking listener-specific localisation behaviour within a full-sphere HRTF framework. Applied to the SONICOM HRTF dataset across 33 listeners, the model reliably recovers its parameters and accurately predicts individual localisation performance, as validated through behavioural experiments. These results establish a statistically rigorous, individually calibrated listener model that opens the door to principled spatial audio personalisation in AR systems, from automatic HRTF selection to perceptual quality evaluation.