Multimodal Learning 相关度: 9/10

SEATrack: Simple, Efficient, and Adaptive Multimodal Tracker

Junbin Su, Ziteng Xue, Shihui Zhang, Kun Chen, Weiming Hu, Zhipeng Zhang
arXiv: 2604.12502v1 发布: 2026-04-14 更新: 2026-04-14

AI 摘要

SEATrack是一种高效多模态跟踪器,通过对齐模态特征和高效融合实现性能与效率的平衡。

主要贡献

  • 提出AMG-LoRA,动态对齐和细化跨模态注意力图。
  • 引入HMoE进行全局关系建模,实现高效跨模态融合。
  • 在RGB-T、RGB-D和RGB-E跟踪任务上实现了性能和效率的提升。

方法论

提出AMG-LoRA对齐注意力图,并使用HMoE进行高效的全局跨模态融合,实现性能和效率的平衡。

原文摘要

Parameter-efficient fine-tuning (PEFT) in multimodal tracking reveals a concerning trend where recent performance gains are often achieved at the cost of inflated parameter budgets, which fundamentally erodes PEFT's efficiency promise. In this work, we introduce SEATrack, a Simple, Efficient, and Adaptive two-stream multimodal tracker that tackles this performance-efficiency dilemma from two complementary perspectives. We first prioritize cross-modal alignment of matching responses, an underexplored yet pivotal factor that we argue is essential for breaking the trade-off. Specifically, we observe that modality-specific biases in existing two-stream methods generate conflicting matching attention maps, thereby hindering effective joint representation learning. To mitigate this, we propose AMG-LoRA, which seamlessly integrates Low-Rank Adaptation (LoRA) for domain adaptation with Adaptive Mutual Guidance (AMG) to dynamically refine and align attention maps across modalities. We then depart from conventional local fusion approaches by introducing a Hierarchical Mixture of Experts (HMoE) that enables efficient global relation modeling, effectively balancing expressiveness and computational efficiency in cross-modal fusion. Equipped with these innovations, SEATrack advances notable progress over state-of-the-art methods in balancing performance with efficiency across RGB-T, RGB-D, and RGB-E tracking tasks. \href{https://github.com/AutoLab-SAI-SJTU/SEATrack}{\textcolor{cyan}{Code is available}}.

标签

多模态跟踪 跨模态对齐 高效融合 视觉跟踪

arXiv 分类

cs.CV cs.AI