AI Agents 相关度: 8/10

Golden Handcuffs make safer AI agents

Aram Ebtekar, Michael K. Cohen
arXiv: 2604.13609v1 发布: 2026-04-15 更新: 2026-04-15

AI 摘要

通过引入负奖励和导师机制,提高强化学习智能体在复杂环境中的安全性和可靠性。

主要贡献

  • 提出一种基于贝叶斯的风险缓解方法,使用负奖励来引导agent
  • 设计了一种导师介入机制,在agent行为风险过高时进行干预
  • 证明了该方法的有效性,包括能力和安全性

方法论

扩展agent的主观奖励范围,并设计一个导师介入机制。通过理论证明来验证方法的有效性。

原文摘要

Reinforcement learners can attain high reward through novel unintended strategies. We study a Bayesian mitigation for general environments: we expand the agent's subjective reward range to include a large negative value $-L$, while the true environment's rewards lie in $[0,1]$. After observing consistently high rewards, the Bayesian policy becomes risk-averse to novel schemes that plausibly lead to $-L$. We design a simple override mechanism that yields control to a safe mentor whenever the predicted value drops below a fixed threshold. We prove two properties of the resulting agent: (i) Capability: using mentor-guided exploration with vanishing frequency, the agent attains sublinear regret against its best mentor. (ii) Safety: no decidable low-complexity predicate is triggered by the optimizing policy before it is triggered by a mentor.

标签

强化学习 安全AI 贝叶斯方法 导师学习

arXiv 分类

cs.LG cs.AI