AI Agents 相关度: 9/10

Stop Fixating on Prompts: Reasoning Hijacking and Constraint Tightening for Red-Teaming LLM Agents

Yanxu Mao, Peipei Liu, Tiehan Cui, Congying Liu, Mingzhe Xing, Datao You
arXiv: 2604.05549v1 发布: 2026-04-07 更新: 2026-04-07

AI 摘要

JailAgent框架通过操纵推理轨迹和记忆检索,实现无需修改prompt的LLM Agent红队测试。

主要贡献

  • 提出 JailAgent 框架
  • 避免修改 prompt 进行红队测试
  • 通过推理劫持和约束收紧提升攻击效果

方法论

JailAgent通过触发词提取、推理劫持和约束收紧三个阶段操纵Agent的推理轨迹和记忆检索。

原文摘要

With the widespread application of LLM-based agents across various domains, their complexity has introduced new security threats. Existing red-team methods mostly rely on modifying user prompts, which lack adaptability to new data and may impact the agent's performance. To address the challenge, this paper proposes the JailAgent framework, which completely avoids modifying the user prompt. Specifically, it implicitly manipulates the agent's reasoning trajectory and memory retrieval with three key stages: Trigger Extraction, Reasoning Hijacking, and Constraint Tightening. Through precise trigger identification, real-time adaptive mechanisms, and an optimized objective function, JailAgent demonstrates outstanding performance in cross-model and cross-scenario environments.

标签

LLM Agent Red-teaming Security Reasoning

arXiv 分类

cs.CL