SkillTrojan: Backdoor Attacks on Skill-Based Agent Systems
AI 摘要
SkillTrojan提出了一种针对技能型代理系统的后门攻击方法,通过恶意技能组合实现攻击。
主要贡献
- 提出SkillTrojan后门攻击方法
- 创建包含3000+后门技能的数据集
- 评估SkillTrojan在代码代理环境中的有效性
方法论
通过在技能中嵌入恶意逻辑,利用技能组合重构和执行攻击者指定的加密payload,并支持自动合成后门技能。
原文摘要
Skill-based agent systems tackle complex tasks by composing reusable skills, improving modularity and scalability while introducing a largely unexamined security attack surface. We propose SkillTrojan, a backdoor attack that targets skill implementations rather than model parameters or training data. SkillTrojan embeds malicious logic inside otherwise plausible skills and leverages standard skill composition to reconstruct and execute an attacker-specified payload. The attack partitions an encrypted payload across multiple benign-looking skill invocations and activates only under a predefined trigger. SkillTrojan also supports automated synthesis of backdoored skills from arbitrary skill templates, enabling scalable propagation across skill-based agent ecosystems. To enable systematic evaluation, we release a dataset of 3,000+ curated backdoored skills spanning diverse skill patterns and trigger-payload configurations. We instantiate SkillTrojan in a representative code-based agent setting and evaluate both clean-task utility and attack success rate. Our results show that skill-level backdoors can be highly effective with minimal degradation of benign behavior, exposing a critical blind spot in current skill-based agent architectures and motivating defenses that explicitly reason about skill composition and execution. Concretely, on EHR SQL, SkillTrojan attains up to 97.2% ASR while maintaining 89.3% clean ACC on GPT-5.2-1211-Global.