Multimodal Learning 相关度: 9/10

POINTS-Seeker: Towards Training a Multimodal Agentic Search Model from Scratch

Yikun Liu, Yuan Liu, Le Tian, Xiao Zhou, Jiangchao Yao, Yanfeng Wang, Weidi Xie
arXiv: 2604.14029v1 发布: 2026-04-15 更新: 2026-04-15

AI 摘要

从零训练多模态Agentic搜索模型POINTS-Seeker,解决长程知识密集型视觉推理难题。

主要贡献

  • Agentic Seeding,为激发Agentic行为奠定基础
  • V-Fold,一种自适应历史感知压缩方案
  • POINTS-Seeker-8B模型,在多项基准测试中超越现有模型

方法论

引入Agentic Seeding和V-Fold压缩方案,从头训练多模态Agentic搜索模型。

原文摘要

While Large Multimodal Models (LMMs) demonstrate impressive visual perception, they remain epistemically constrained by their static parametric knowledge. To transcend these boundaries, multimodal search models have been adopted to actively interact with the external environment for evidence retrieval. Diverging from prevailing paradigms that merely retrofit general LMMs with search tools as modular extensions, we explore the potential of building a multimodal agentic search model from scratch. Specifically, we make the following contributions: (i) we introduce Agentic Seeding, a dedicated phase designed to weave the foundational precursors necessary for eliciting agentic behaviors; (ii) we uncover a performance bottleneck in long-horizon interactions, where the increasing volume of interaction history overwhelms the model's ability to locate ground-truth evidence. To mitigate this, we propose V-Fold, an adaptive history-aware compression scheme that preserves recent dialogue turns in high fidelity while folding historical context into the visual space via rendering; and (iii) we develop POINTS-Seeker-8B, a state-of-the-art multimodal agentic search model that consistently outperforms existing models across six diverse benchmarks, effectively resolving the challenges of long-horizon, knowledge-intensive visual reasoning.

标签

多模态学习 AI Agent 视觉推理 模型训练 知识检索

arXiv 分类

cs.CV