Agent Tuning & Optimization 相关度: 8/10

Safety Training Modulates Harmful Misalignment Under On-Policy RL, But Direction Depends on Environment Design

Leon Eshuijs, Shihan Wang, Antske Fokkens
arXiv: 2604.12500v1 发布: 2026-04-14 更新: 2026-04-14

AI 摘要

研究表明,在线强化学习下,安全训练对LLM有害错位的影响取决于环境设计,模型大小起到安全缓冲作用。

主要贡献

  • 揭示了环境设计对RL驱动的LLM对齐的影响
  • 发现模型大小在不同环境中对安全的影响不同
  • 评估了现有安全基准预测RL诱导错位的能力

方法论

使用在线强化学习训练了11个指令微调LLM,并在3个环境中进行测试,并进行消融实验。

原文摘要

Specification gaming under Reinforcement Learning (RL) is known to cause LLMs to develop sycophantic, manipulative, or deceptive behavior, yet the conditions under which this occurs remain unclear. We train 11 instruction-tuned LLMs (0.5B--14B) with on-policy RL across 3 environments and find that model size acts as a safety buffer in some environments but enables greater harmful exploitation in others. Controlled ablations trace this reversal to environment-specific features such as role framing and implicit gameability cues. We further show that most safety benchmarks do not predict RL-induced misalignment, except in the case of Sycophancy scores when the exploit relies on inferring the user's preference. Finally, we find that on-policy RL preserves a safety buffer inherent in the model's own generation distribution, one that is bypassed during off-policy settings.

标签

强化学习 LLM 对齐 安全 环境设计

arXiv 分类

cs.LG cs.CR