Multimodal Learning 相关度: 9/10

Doc-V*:Coarse-to-Fine Interactive Visual Reasoning for Multi-Page Document VQA

Yuanlei Zheng, Pei Fu, Hang Li, Ziyang Wang, Yuyi Zhang, Wenyu Ruan, Xiaojin Zhang, Zhongyu Wei, Zhenbo Luo, Jian Luan, Wei Chen, Xiang Bai
arXiv: 2604.13731v1 发布: 2026-04-15 更新: 2026-04-15

AI 摘要

Doc-V*是一个OCR-free的agent框架,通过主动导航和证据聚合解决多页文档VQA问题。

主要贡献

  • 提出Doc-V* agent框架
  • 基于模仿学习和群体相对策略优化进行训练
  • 在多个基准测试中表现优于现有方法

方法论

构建OCR-free agent,通过语义检索和页面获取进行主动导航,在结构化工作记忆中聚合证据,实现多页文档的VQA。

原文摘要

Multi-page Document Visual Question Answering requires reasoning over semantics, layouts, and visual elements in long, visually dense documents. Existing OCR-free methods face a trade-off between capacity and precision: end-to-end models scale poorly with document length, while visual retrieval-based pipelines are brittle and passive. We propose Doc-$V^*$, an \textbf{OCR-free agentic} framework that casts multi-page DocVQA as sequential evidence aggregation. Doc-$V^*$ begins with a thumbnail overview, then actively navigates via semantic retrieval and targeted page fetching, and aggregates evidence in a structured working memory for grounded reasoning. Trained by imitation learning from expert trajectories and further optimized with Group Relative Policy Optimization, Doc-$V^*$ balances answer accuracy with evidence-seeking efficiency. Across five benchmarks, Doc-$V^*$ outperforms open-source baselines and approaches proprietary models, improving out-of-domain performance by up to \textbf{47.9\%} over RAG baseline. Other results reveal effective evidence aggregation with selective attention, not increased input pages.

标签

DocVQA Multi-page Document Agent Visual Reasoning

arXiv 分类

cs.CL