AI Agents 相关度: 9/10

Beyond State Consistency: Behavior Consistency in Text-Based World Models

Youling Huang, Guanqiao Chen, Junchi Yao, Lu Wang, Fangkai Yang, Chao Du, ChenZhuo Zhao, Pu Zhao, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang
arXiv: 2604.13824v1 发布: 2026-04-15 更新: 2026-04-15

AI 摘要

提出了一种新的行为一致性训练范式,通过优化行为一致性奖励(BehR)来提升文本世界模型与真实环境的功能一致性。

主要贡献

  • 提出了行为一致性奖励(BehR)作为步级评估指标。
  • 设计了基于BehR的训练范式,用于优化文本世界模型。
  • 实验证明BehR训练可以提升长期对齐,并改善离线评估和在线规划效果。

方法论

通过参考agent,衡量真实状态和模型预测状态下动作可能性变化,以此设计BehR作为奖励函数,优化世界模型。

原文摘要

World models have been emerging as critical components for assessing the consequences of actions generated by interactive agents in online planning and offline evaluation. In text-based environments, world models are typically evaluated and trained with single-step metrics such as Exact Match, aiming to improve the similarity between predicted and real-world states, but such metrics have been shown to be insufficient for capturing actual agent behavior. To address this issue, we introduce a new behavior-aligned training paradigm aimed at improving the functional consistency between the world model and the real environment. This paradigm focuses on optimizing a tractable step-level metric named Behavior Consistency Reward (BehR), which measures how much the likelihood of a logged next action changes between the real state and the world-model-predicted state under a frozen Reference Agent. Experiments on WebShop and TextWorld show that BehR-based training improves long-term alignment in several settings, with the clearest gains in WebShop and less movement in near-ceiling regimes, while preserving or improving single-step prediction quality in three of four settings. World models trained with BehR also achieve lower false positives in offline surrogate evaluation and show modest but encouraging gains in inference-time lookahead planning.

标签

world model behavior consistency text-based environment reinforcement learning

arXiv 分类

cs.LG