Speaker
Description
This paper presents SPSC-HCM-16C, a benchmarkdataset1for patient-level respiratory disease classification from synchronized 16-channel lung soundrecordings collected in a real clinical environment.The dataset contains recordings from 183 subjects,including healthy controls and four respiratory diseasegroups, and supports three diagnostic settings: 2-class,3-class, and 5-class classification. The aim is to establish a standardized and reproducible benchmark forfuture research in multi-channel computational auscultation. To this end, we define subject-independentevaluation protocols, including a fixed 60/20/20 splitand an 80/20 split with 5-fold cross-validation on thetraining-validation subset, together with a transparent reference baseline. The baseline uses MFCC features, a modified 16-channel ResNet101 backbone withpartial fine-tuning of Layer3, and mean pooling forpatient-level aggregation. Under the fixed 60/20/20split, the baseline achieves Macro-F1 scores of 93.68%,77.06%, and 44.61% for the 2-class, 3-class, and 5-classtasks, respectively, with consistent trends observedunder 5-fold cross-validation. These results establish areproducible reference for future comparison and highlight the challenges posed by weak supervision, classimbalance, and inter-class similarity in multi-channellung sound analysis.