Securing Retrieval-Augmented Generation: A Taxonomy of Attacks, Defenses, and Future Directions
AI 摘要
该论文系统性地分析了RAG中的安全威胁,提出了攻击、防御和未来方向的分类框架。
主要贡献
- 定义了安全RAG的边界,区分了LLM固有风险与RAG引入的风险
- 提出了基于RAG工作流阶段的安全威胁分类体系
- 总结了现有攻击、防御和评估基准,并指出了当前防御的局限性
方法论
通过文献综述,将RAG工作流抽象为六个阶段,并基于信任边界和安全表面组织文献。
原文摘要
Retrieval-augmented generation (RAG) significantly enhances large language models (LLMs) but introduces novel security risks through external knowledge access. While existing studies cover various RAG vulnerabilities, they often conflate inherent LLM risks with those specifically introduced by RAG. In this paper, we propose that secure RAG is fundamentally about the security of the external knowledge-access pipeline. We establish an operational boundary to separate inherent LLM flaws from RAG-introduced or RAG-amplified threats. Guided by this perspective, we abstract the RAG workflow into six stages and organize the literature around three trust boundaries and four primary security surfaces, including pre-retrieval knowledge corruption, retrieval-time access manipulation, downstream context exploitation, and knowledge exfiltration. By systematically reviewing the corresponding attacks, defenses, remediation mechanisms, and evaluation benchmarks, we reveal that current defenses remain largely reactive and fragmented. Finally, we discuss these gaps and highlight future directions toward layered, boundary-aware protection across the entire knowledge-access lifecycle.