Multimodal Learning 相关度: 9/10

Learning Shared Sentiment Prototypes for Adaptive Multimodal Sentiment Analysis

Chen Su, Yuanhe Tian, Yan Song
arXiv: 2604.05873v1 发布: 2026-04-07 更新: 2026-04-07

AI 摘要

PRISM通过共享原型空间进行多模态情感分析,并动态调整模态权重,提升情感预测精度。

主要贡献

  • 提出基于共享原型空间的多模态情感分析框架PRISM
  • 动态模态重加权机制,允许在推理过程中持续优化模态贡献
  • 在三个基准数据集上验证了PRISM的优越性

方法论

构建共享原型空间进行跨模态比较和自适应融合,利用动态模态重加权进行情感推理。

原文摘要

Multimodal sentiment analysis (MSA) aims to predict human sentiment from textual, acoustic, and visual information in videos. Recent studies improve multimodal fusion by modeling modality interaction and assigning different modality weights. However, they usually compress diverse sentiment cues into a single compact representation before sentiment reasoning. This early aggregation makes it difficult to preserve the internal structure of sentiment evidence, where different cues may complement, conflict with, or differ in reliability from each other. In addition, modality importance is often determined only once during fusion, so later reasoning cannot further adjust modality contributions. To address these issues, we propose PRISM, a framework that unifies structured affective extraction and adaptive modality evaluation. PRISM organizes multimodal evidence in a shared prototype space, which supports structured cross-modal comparison and adaptive fusion. It further applies dynamic modality reweighting during reasoning, allowing modality contributions to be continuously refined as semantic interactions become deeper. Experiments on three benchmark datasets show that PRISM outperforms representative baselines.

标签

多模态情感分析 模态融合 共享原型 动态权重

arXiv 分类

cs.MM