Benchmarking Open-Source Speech Emotion Recognition in Naturalistic Mandarin Spine Clinic Consultations: A Pilot Validation Study
Speech emotion recognition (SER) may enable passive affect monitoring in clinical encounters, but most systems are validated on acted laboratory speech rather than naturalistic Mandarin outpatient consultations. We benchmarked three open-source SER models (emotion2vec+, SenseVoice, FunASR) against a researcher-consensus reference in naturalistic spine-clinic speech, assessing minority-state detection under class imbalance. In a retrospective analysis of prospectively collected single-center recordings, audio was loudness-normalized and only conversations among patients, family members, and clinicians were retained. Sixty-five utterances (5-50 s; one per participant; 31 patients, 34 family members) were labeled by six calibrated annotators into six categories (Happy, Sad, Fear, Anger, Neutral, Surprised). Consensus used majority vote with Fleiss kappa filtering and clinician adjudication for low-agreement segments. Metrics included unweighted accuracy (UA), macro-average per-class accuracy, class- and sample-level weighted accuracy (WA), and F1 with bootstrap 95% CIs. Labels were imbalanced (Neutral 58.5%); median Fleiss kappa was 0.230 (IQR 0.134-0.519). SenseVoice and FunASR achieved UA 61.5% (95% CI 49.2-73.8%), macro-average per-class accuracy 87.2%, class-level WA 92.5%, and sample-level WA 24.6%. emotion2vec+ yielded UA 55.4% (95% CI 42.5-67.7%) and macro-average per-class accuracy 85.1%, with class-level WA 90.6% and sample-level WA 24.2%. Despite high inter-model agreement (90.8%), all models had near-zero recall for Sad, Fear, Anger, and Surprised. In this pilot, majority-class accuracy was misleading: SER poorly detected minority emotions against a noisy naturalistic reference. Clinical deployment readiness cannot be inferred from acted-corpus benchmarks without domain adaptation, multimodal modeling, stronger reference standards, and outcome validation.