Agent Tuning & Optimization 相关度: 8/10

Route to Rome Attack: Directing LLM Routers to Expensive Models via Adversarial Suffix Optimization

Haochun Tang, Yuliang Yan, Jiahua Lu, Huaxiao Liu, Enyan Dai
arXiv: 2604.15022v1 发布: 2026-04-16 更新: 2026-04-16

AI 摘要

R$^2$A攻击通过优化对抗后缀,诱导黑盒LLM路由器选择高成本模型。

主要贡献

  • 提出了一种新的黑盒LLM路由攻击方法R$^2$A
  • 使用混合集成代理路由器模拟黑盒路由器
  • 设计了一种后缀优化算法,针对集成的代理路由器

方法论

使用混合集成代理模型模拟黑盒路由,并通过后缀优化算法生成对抗样本。

原文摘要

Cost-aware routing dynamically dispatches user queries to models of varying capability to balance performance and inference cost. However, the routing strategy introduces a new security concern that adversaries may manipulate the router to consistently select expensive high-capability models. Existing routing attacks depend on either white-box access or heuristic prompts, rendering them ineffective in real-world black-box scenarios. In this work, we propose R$^2$A, which aims to mislead black-box LLM routers to expensive models via adversarial suffix optimization. Specifically, R$^2$A deploys a hybrid ensemble surrogate router to mimic the black-box router. A suffix optimization algorithm is further adapted for the ensemble-based surrogate. Extensive experiments on multiple open-source and commercial routing systems demonstrate that {R$^2$A} significantly increases the routing rate to expensive models on queries of different distributions. Code and examples: https://github.com/thcxiker/R2A-Attack.

标签

LLM 路由 安全 对抗攻击 优化

arXiv 分类

cs.CR cs.AI cs.CL cs.LG