Multimodal Learning 相关度: 9/10

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs

Haicheng Wang, Yuan Liu, Yikun Liu, Zhemeng Yu, Zhongyin Zhao, Yangxiu You, Zilin Yu, Le Tian, Xiao Zhou, Jie Zhou, Weidi Xie, Yanfeng Wang
arXiv: 2604.11627v1 发布: 2026-04-13 更新: 2026-04-13

AI 摘要

POINTS-Long提出了一种双模态MLLM,通过动态视觉token缩放,提升长视频理解效率。

主要贡献

  • 提出了一种双模态MLLM结构,包含focus和standby两种模式
  • 设计了动态视觉token缩放机制,在效率和精度之间进行权衡
  • 实现了流式视觉理解,支持超长视觉记忆

方法论

设计双模态视觉感知,focus模式用于精细任务,standby模式用于通用任务,并采用动态可分离KV-cache支持流式处理。

原文摘要

Multimodal Large Language Models (MLLMs) have recently demonstrated remarkable capabilities in cross-modal understanding and generation. However, the rapid growth of visual token sequences--especially in long-video and streaming scenarios--poses a major challenge to their scalability and real-world deployment. Thus, we introduce POINTS-Long, a native dual-mode MLLM featuring dynamic visual token scaling inspired by the human visual system. The model supports two complementary perception modes: focus mode and standby mode, enabling users to dynamically trade off efficiency and accuracy during inference. On fine-grained visual tasks, the focus mode retains the optimal performance, while on long-form general visual understanding, the standby mode retains 97.7-99.7% of the original accuracy using only 1/40-1/10th of the visual tokens. Moreover, POINTS-Long natively supports streaming visual understanding via a dynamically detachable KV-cache design, allowing efficient maintenance of ultra-long visual memory. Our work provides new insights into the design of future MLLMs and lays the foundation for adaptive and efficient long-form visual understanding.

标签

MLLM 长视频理解 视觉token缩放 流式处理

arXiv 分类

cs.CV