Multimodal Learning 相关度: 9/10

Mixture-of-Modality-Experts with Holistic Token Learning for Fine-Grained Multimodal Visual Analytics in Driver Action Recognition

Tianyi Liu, Yiming Li, Wenqian Wang, Jiaojiao Wang, Chen Cai, Yi Wang, Kim-Hui Yap
arXiv: 2604.05947v1 发布: 2026-04-07 更新: 2026-04-07

AI 摘要

提出MoME和HTL框架,通过模态专家混合和整体token学习提升驾驶员行为识别的细粒度多模态理解能力。

主要贡献

  • 提出 Mixture-of-Modality-Experts (MoME) 框架
  • 设计 Holistic Token Learning (HTL) 策略
  • 在驾驶员行为识别任务上验证框架的有效性

方法论

MoME实现模态专家自适应协作,HTL通过类token和时空token增强专家内部优化和跨专家知识迁移,提升多模态融合效果。

原文摘要

Robust multimodal visual analytics remains challenging when heterogeneous modalities provide complementary but input-dependent evidence for decision-making.Existing multimodal learning methods mainly rely on fixed fusion modules or predefined cross-modal interactions, which are often insufficient to adapt to changing modality reliability and to capture fine-grained action cues. To address this issue, we propose a Mixture-of-Modality-Experts (MoME) framework with a Holistic Token Learning (HTL) strategy. MoME enables adaptive collaboration among modality-specific experts, while HTL improves both intra-expert refinement and inter-expert knowledge transfer through class tokens and spatio-temporal tokens. In this way, our method forms a knowledge-centric multimodal learning framework that improves expert specialization while reducing ambiguity in multimodal fusion.We validate the proposed framework on driver action recognition as a representative multimodal understanding taskThe experimental results on the public benchmark show that the proposed MoME framework and the HTL strategy jointly outperform representative single-modal and multimodal baselines. Additional ablation, validation, and visualization results further verify that the proposed HTL strategy improves subtle multimodal understanding and offers better interpretability.

标签

多模态学习 驾驶员行为识别 专家混合模型 Token学习

arXiv 分类

cs.CV