Multimodal Learning 相关度: 9/10

DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding

Hao Yan, Yuliang Liu, Xingchen Liu, Yuyi Zhang, Minghui Liao, Jihao Wu, Wei Chen, Xiang Bai
arXiv: 2604.12812v1 发布: 2026-04-14 更新: 2026-04-14

AI 摘要

DocSeeker通过结构化流程和证据引导提升MLLM在长文档理解中的表现,解决了信噪比低和监督稀疏问题。

主要贡献

  • 提出了结构化的分析、定位和推理工作流程
  • 设计了两阶段训练框架(监督微调+证据感知相对策略优化)
  • 引入了证据引导的分辨率分配策略

方法论

采用知识蒸馏生成高质量数据,进行监督微调,并使用证据感知的相对策略优化方法联合优化证据定位和答案准确性。

原文摘要

Existing Multimodal Large Language Models (MLLMs) suffer from significant performance degradation on the long document understanding task as document length increases. This stems from two fundamental challenges: 1) a low Signal-to-Noise Ratio (SNR), with crucial evidence buried in irrelevant pages; and 2) supervision scarcity, as datasets offering only final short answers provide a weak learning signal. In this paper, we address these challenges by proposing a paradigm that requires the model to execute a structured ``\textbf{Analysis}, \textbf{Localization} and \textbf{Reasoning}'' workflow. To instill this capability, we design a two-stage training framework: we first perform Supervised Fine-Tuning on high-quality data generated via an efficient knowledge distillation strategy. Subsequently, we employ an Evidence-aware Group Relative Policy Optimization which jointly optimizes for both evidence localization and answer accuracy. Additionally, we introduce a Evidence-Guided Resolution Allocation strategy to mitigate memory constraints of training on multi-pages documents. Extensive experiments demonstrate that DocSeeker achieves superior performance on both in-domain and out-of-domain tasks. We show it robustly generalizes from short-page training to ultra-long documents and is naturally synergistic with visual Retrieval-Augmented Generation systems, serving as a solid foundation for their implementation.

标签

Multimodal Learning Long Document Understanding Visual Reasoning

arXiv 分类

cs.AI