Multimodal Learning 相关度: 9/10

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs

Muhammad Kamran Janjua, Hugo Silva, Di Niu, Bahador Rashidi
arXiv: 2604.12896v1 发布: 2026-04-14 更新: 2026-04-14

AI 摘要

P$^2$通过语言化视觉工具输出,显著提升多模态大模型视觉推理能力。

主要贡献

  • 提出Perception Programs (P$^2$)方法,无需训练即可提升MLLM的视觉工具推理能力
  • P$^2$将视觉工具输出转换为紧凑、结构化的语言描述
  • 实验证明P$^2$在多个perception-centric任务上显著优于现有方法

方法论

P$^2$是一种training-free方法,通过将视觉工具输出重写为语言形式,使MLLM能够直接解析和推理。

原文摘要

Multimodal language models (MLLMs) are increasingly paired with vision tools (e.g., depth, flow, correspondence) to enhance visual reasoning. However, despite access to these tool-generated visual cues, MLLMs often fail to benefit from them. Existing approaches typically feed raw tool outputs into the model, but these dense, pixel-level representations are misaligned with the language-native reasoning strengths of LLMs, leading to weak perception and reliance on language priors. We argue that, in problems where vision tools can provide the necessary visual cues, the bottleneck is not more tool calls or larger MLLMs, it is how tool outputs are represented. We introduce Perception Programs (P$^2$), a training-free, model-agnostic method that rewrites tool outputs into compact, structured, language-native summaries that MLLMs can directly parse and reason over. Across six perception-centric tasks in BLINK, P$^2$ consistently yields large improvements over base models and raw tool-augmented baselines. With GPT-5 Mini as the base model, P$^2$ raises its accuracy from 41.35\% to 86.47\% on multi-view reasoning, from 52.42\% to 81.45\% on relative depth, and achieves a 22\% average gain across tasks, setting new state-of-the-art results. Even on smaller MLLMs, e.g., InternVL3.5-4B and Qwen3VL-4B, we observe 15-40\% absolute gains from P$^2$, surpassing prior agentic, supervised, and RL-based tool-use methods-without any training or model modifications.

标签

MLLM 视觉推理 工具使用 感知程序 语言化视觉

arXiv 分类

cs.CV cs.LG