Back to Search View Original Cite This Article

Abstract

<title>Abstract</title> <p>Major Depressive Disorder (MDD) remains substantially underdiagnosed as current clinical assessment depends on subjective judgement and resource-intensive interviews. Automated multimodal analysis of clinical interviews offers a promising path toward objective, scalable screening, yet two obstacles limit progress: existing approaches rely on static fusion strategies that ignore per subject modality reliability, and a considerable share of reported results are inflated by evaluation leakage that permits test-set information to influence model selection. This paper proposes a three-modality framework utilizing transcript text embeddings, acoustic features extracted by WavLM, and a third modality derived from Speech Emotion Captioning. A large language model produces free-form natural language descriptions of emotional state directly from raw audio. We introduce Confidence-Gated Feature Modulation (CGFM), a subject-adaptive, parameter-free mechanism that weights each modality by its predictive confidence per subject, dynamically suppressing uninformative inputs. Under a nested cross-validation protocol applied uniformly to all baselines, the proposed method achieves a Macro F1 of 0.9125 on E-DAIC, and generalises to DAIC-WOZ (F1 = 0.6438) and CMDC (F1 = 0.9369). Ablation studies confirm that all three modalities contribute to the model performance.</p>

Show More

Keywords

modality model clinical interviews subject

Related Articles

PORE

About

Connect