AI Agents 相关度: 10/10

ClawGuard: A Runtime Security Framework for Tool-Augmented LLM Agents Against Indirect Prompt Injection

Wei Zhao, Zhe Li, Peixin Zhang, Jun Sun
arXiv: 2604.11790v1 发布: 2026-04-13 更新: 2026-04-13

AI 摘要

ClawGuard通过在工具调用边界实施用户确认规则,防御工具增强型LLM代理的间接prompt注入攻击。

主要贡献

  • 提出ClawGuard框架,实现工具调用边界的确定性安全策略
  • 无需模型修改或架构变更,有效防御多种注入攻击
  • 实验验证了ClawGuard在不影响代理效用的前提下的安全性能

方法论

ClawGuard通过在每次工具调用前,根据用户目标自动推导访问约束,拦截恶意工具调用,实现运行时安全。

原文摘要

Tool-augmented Large Language Model (LLM) agents have demonstrated impressive capabilities in automating complex, multi-step real-world tasks, yet remain vulnerable to indirect prompt injection. Adversaries exploit this weakness by embedding malicious instructions within tool-returned content, which agents directly incorporate into their conversation history as trusted observations. This vulnerability manifests across three primary attack channels: web and local content injection, MCP server injection, and skill file injection. To address these vulnerabilities, we introduce \textsc{ClawGuard}, a novel runtime security framework that enforces a user-confirmed rule set at every tool-call boundary, transforming unreliable alignment-dependent defense into a deterministic, auditable mechanism that intercepts adversarial tool calls before any real-world effect is produced. By automatically deriving task-specific access constraints from the user's stated objective prior to any external tool invocation, \textsc{ClawGuard} blocks all three injection pathways without model modification or infrastructure change. Experiments across five state-of-the-art language models on AgentDojo, SkillInject, and MCPSafeBench demonstrate that \textsc{ClawGuard} achieves robust protection against indirect prompt injection without compromising agent utility. This work establishes deterministic tool-call boundary enforcement as an effective defense mechanism for secure agentic AI systems, requiring neither safety-specific fine-tuning nor architectural modification. Code is publicly available at https://github.com/Claw-Guard/ClawGuard.

标签

AI Agents Security Prompt Injection Tool Use

arXiv 分类

cs.CR cs.AI