Speaker
Description
Human spatial hearing relies on the integration of multiple binaural cues, including interaural time differences (ITDs), interaural level differences (ILDs), and spectral information. While psychophysical experiments have characterized how these cues are combined across a wide range of conditions, linking perceptual behavior to underlying neural mechanisms remains a central challenge. Deep neural networks (DNNs) trained on sound localization tasks provide a promising framework for bridging this gap.Here, we evaluate a DNN model of spatial hearing with a focus on its ability to reproduce human psychophysical data across diverse and often highly artificial stimulus conditions. The model captures key aspects of human perception, including envelope ITD–based lateralization, interactions between ITD and ILD, robustness to strong reflections, and stimulus-dependent spectral cue weighting. While these properties were not directly imposed by design, input characteristics shaped by cochlear filtering and hair-cell nonlinearities constrain the functional space of both artificial and biological binaural processing.At the same time, systematic deviations from human performance are observed. The model fails in conditions that rely on weak or ambiguous cues, such as broadband noise with supranatural ITDs, and does not reproduce known perceptual biases in pure-tone localization. These discrepancies point to differences in cue integration strategies and suggest missing biological constraints, such as frequency-dependent priors or limitations in combining binaural information.By directly comparing model behavior with established psychophysical paradigms, this work demonstrates how DNNs can be used not only to reproduce but also to interrogate human spatial hearing, providing a framework for generating targeted experimental predictions and refining theories of binaural perception.