Multimodal Learning 相关度: 9/10

RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time

Haozhe Wang, Cong Wei, Weiming Ren, Jiaming Liu, Fangzhen Lin, Wenhu Chen
arXiv: 2604.11626v1 发布: 2026-04-13 更新: 2026-04-13

AI 摘要

提出RationalRewards模型,通过推理奖励提升视觉生成效果,无需参数更新即可优化生成结果。

主要贡献

  • 提出RationalRewards模型,使用多维度评判提升视觉生成效果
  • 提出Preference-Anchored Rationalization (PARROT)框架,利用偏好数据生成高质量推理
  • 测试时利用Critique-Refine循环改进生成结果,无需参数更新

方法论

训练奖励模型生成显式、多维度的评判,通过强化学习在训练时提供细粒度奖励,测试时通过生成-评判-改进循环优化。

原文摘要

Most reward models for visual generation reduce rich human judgments to a single unexplained score, discarding the reasoning that underlies preference. We show that teaching reward models to produce explicit, multi-dimensional critiques before scoring transforms them from passive evaluators into active optimization tools, improving generators in two complementary ways: at training time, structured rationales provide interpretable, fine-grained rewards for reinforcement learning; at test time, a Generate-Critique-Refine loop turns critiques into targeted prompt revisions that improve outputs without any parameter updates. To train such a reward model without costly rationale annotations, we introduce Preference-Anchored Rationalization (PARROT), a principled framework that recovers high-quality rationales from readily available preference data through anchored generation, consistency filtering, and distillation. The resulting model, RationalRewards (8B), achieves state-of-the-art preference prediction among open-source reward models, competitive with Gemini-2.5-Pro, while using 10-20x less training data than comparable baselines. As an RL reward, it consistently improves text-to-image and image-editing generators beyond scalar alternatives. Most strikingly, its test-time critique-and-refine loop matches or exceeds RL-based fine-tuning on several benchmarks, suggesting that structured reasoning can unlock latent capabilities in existing generators that suboptimal prompts fail to elicit.

标签

视觉生成 奖励模型 强化学习 推理 文本到图像

arXiv 分类

cs.AI cs.LG