Golden Handcuffs make safer AI agents
AI 摘要
通过引入负奖励和导师机制,提高强化学习智能体在复杂环境中的安全性和可靠性。
主要贡献
- 提出一种基于贝叶斯的风险缓解方法,使用负奖励来引导agent
- 设计了一种导师介入机制,在agent行为风险过高时进行干预
- 证明了该方法的有效性,包括能力和安全性
方法论
扩展agent的主观奖励范围,并设计一个导师介入机制。通过理论证明来验证方法的有效性。
原文摘要
Reinforcement learners can attain high reward through novel unintended strategies. We study a Bayesian mitigation for general environments: we expand the agent's subjective reward range to include a large negative value $-L$, while the true environment's rewards lie in $[0,1]$. After observing consistently high rewards, the Bayesian policy becomes risk-averse to novel schemes that plausibly lead to $-L$. We design a simple override mechanism that yields control to a safe mentor whenever the predicted value drops below a fixed threshold. We prove two properties of the resulting agent: (i) Capability: using mentor-guided exploration with vanishing frequency, the agent attains sublinear regret against its best mentor. (ii) Safety: no decidable low-complexity predicate is triggered by the optimizing policy before it is triggered by a mentor.