Dynamic Summary Generation for Interpretable Multimodal Depression Detection
AI 摘要
提出一种多阶段框架,利用LLM生成动态摘要,提升多模态抑郁症检测的准确性和可解释性。
主要贡献
- 提出基于LLM的动态摘要生成方法
- 构建多阶段抑郁症检测框架
- 融合文本、音频和视频特征提升检测效果
方法论
采用粗到细的多阶段流程,利用LLM生成临床摘要,指导多模态融合模块进行预测,并生成人类可读的评估报告。
原文摘要
Depression remains widely underdiagnosed and undertreated because stigma and subjective symptom ratings hinder reliable screening. To address this challenge, we propose a coarse-to-fine, multi-stage framework that leverages large language models (LLMs) for accurate and interpretable detection. The pipeline performs binary screening, five-class severity classification, and continuous regression. At each stage, an LLM produces progressively richer clinical summaries that guide a multimodal fusion module integrating text, audio, and video features, yielding predictions with transparent rationale. The system then consolidates all summaries into a concise, human-readable assessment report. Experiments on the E-DAIC and CMDC datasets show significant improvements over state-of-the-art baselines in both accuracy and interpretability.