Speaker
Description
We investigate whether the frame-theoretic properties of filterbanks can predict downstream neural network classification performance. Using a dataset of 3,080 African savanna elephant rumble vocalisations classified into four age groups, we train a convolutional neural network on magnitude spectrograms derived from 15 filterbank configurations spanning equivalent rectangular bandwidth (ERB), constant-Q, gammatone, and mel families, including their tight-frame variants.Two findings emerge: First, the condition number kappa = B/A is not a statistically significant predictor of CNN performance within the controlled ERB family (r = 0.738, p = 0.155, n = 5). Second, the considered invalid-frame configurations significantly outperform the valid-frame configurations (p = 0.006). An apparent overall correlation between condition number and F1 of r = 0.954 is identified as a filter-family confound. In summary, we find that in our setting tightness does not benefit CNN classification performance but that the filter family design - specifically frequency resolution and spacing - is the dominant factor.