Multimodal Learning 相关度: 10/10

OpenVLThinkerV2: A Generalist Multimodal Reasoning Model for Multi-domain Visual Tasks

Wenbo Hu, Xin Chen, Yan Gao-Tian, Yihe Deng, Nanyun Peng, Kai-Wei Chang
arXiv: 2604.08539v1 发布: 2026-04-09 更新: 2026-04-09

AI 摘要

OpenVLThinkerV2是一个通用的多模态推理模型,通过改进的强化学习方法,在多领域视觉任务上表现出色。

主要贡献

  • 提出了Gaussian GRPO (G$^2$RPO) 强化学习目标,增强训练稳定性。
  • 引入了响应长度和熵整形机制,平衡感知和推理能力。
  • 发布了OpenVLThinkerV2模型,并在多个基准测试中优于现有模型。

方法论

使用Gaussian GRPO进行强化学习训练,通过响应长度和熵整形机制平衡感知和推理能力,构建通用多模态模型。

原文摘要

Group Relative Policy Optimization (GRPO) has emerged as the de facto Reinforcement Learning (RL) objective driving recent advancements in Multimodal Large Language Models. However, extending this success to open-source multimodal generalist models remains heavily constrained by two primary challenges: the extreme variance in reward topologies across diverse visual tasks, and the inherent difficulty of balancing fine-grained perception with multi-step reasoning capabilities. To address these issues, we introduce Gaussian GRPO (G$^2$RPO), a novel RL training objective that replaces standard linear scaling with non-linear distributional matching. By mathematically forcing the advantage distribution of any given task to strictly converge to a standard normal distribution, $\mathcal{N}(0,1)$, G$^2$RPO theoretically ensures inter-task gradient equity, mitigates vulnerabilities to heavy-tail outliers, and offers symmetric update for positive and negative rewards. Leveraging the enhanced training stability provided by G$^2$RPO, we introduce two task-level shaping mechanisms to seamlessly balance perception and reasoning. First, response length shaping dynamically elicits extended reasoning chains for complex queries while enforce direct outputs to bolster visual grounding. Second, entropy shaping tightly bounds the model's exploration zone, effectively preventing both entropy collapse and entropy explosion. Integrating these methodologies, we present OpenVLThinkerV2, a highly robust, general-purpose multimodal model. Extensive evaluations across 18 diverse benchmarks demonstrate its superior performance over strong open-source and leading proprietary frontier models.

标签

Multimodal Learning Reinforcement Learning Visual Reasoning

arXiv 分类

cs.CV cs.AI cs.CL