AI Agents 相关度: 9/10

Playing Along: Learning a Double-Agent Defender for Belief Steering via Theory of Mind

Hanqi Xiao, Vaidehi Patil, Zaid Khan, Hyunji Lee, Elias Stengel-Eskin, Mohit Bansal
arXiv: 2604.11666v1 发布: 2026-04-13 更新: 2026-04-13

AI 摘要

提出ToM-SB挑战,训练AI Double Agent,提升LLM在对抗场景中的Theory of Mind能力。

主要贡献

  • 提出了一个隐私主题的ToM挑战:ToM-SB。
  • 训练了能够扮演AI Double Agent的模型,并通过强化学习提升其ToM和欺骗能力。
  • 证明了ToM和欺骗成功之间存在双向促进关系,并且结合ToM和欺骗奖励能获得最佳性能。

方法论

通过强化学习训练AI Double Agent,分别使用欺骗奖励和ToM奖励,并评估其在ToM-SB上的表现。

原文摘要

As large language models (LLMs) become the engine behind conversational systems, their ability to reason about the intentions and states of their dialogue partners (i.e., form and use a theory-of-mind, or ToM) becomes increasingly critical for safe interaction with potentially adversarial partners. We propose a novel privacy-themed ToM challenge, ToM for Steering Beliefs (ToM-SB), in which a defender must act as a Double Agent to steer the beliefs of an attacker with partial prior knowledge within a shared universe. To succeed on ToM-SB, the defender must engage with and form a ToM of the attacker, with a goal of fooling the attacker into believing they have succeeded in extracting sensitive information. We find that strong frontier models like Gemini3-Pro and GPT-5.4 struggle on ToM-SB, often failing to fool attackers in hard scenarios with partial attacker prior knowledge, even when prompted to reason about the attacker's beliefs (ToM prompting). To close this gap, we train models on ToM-SB to act as AI Double Agents using reinforcement learning, testing both fooling and ToM rewards. Notably, we find a bidirectionally emergent relationship between ToM and attacker-fooling: rewarding fooling success alone improves ToM, and rewarding ToM alone improves fooling. Across four attackers with different strengths, six defender methods, and both in-distribution and out-of-distribution (OOD) evaluation, we find that gains in ToM and attacker-fooling are well-correlated, highlighting belief modeling as a key driver of success on ToM-SB. AI Double Agents that combine both ToM and fooling rewards yield the strongest fooling and ToM performance, outperforming Gemini3-Pro and GPT-5.4 with ToM prompting on hard scenarios. We also show that ToM-SB and AI Double Agents can be extended to stronger attackers, demonstrating generalization to OOD settings and the upgradability of our task.

标签

LLM Theory of Mind Reinforcement Learning Adversarial Learning Double Agent

arXiv 分类

cs.CL cs.AI cs.LG