NaviRAG: Towards Active Knowledge Navigation for Retrieval-Augmented Generation
AI 摘要
NaviRAG通过知识分层和主动导航提升RAG性能,实现多粒度信息检索和动态规划。
主要贡献
- 提出 NaviRAG 框架,实现主动知识导航
- 利用知识分层结构进行多粒度信息检索
- 通过实验证明 NaviRAG 在长文档 QA 上的优势
方法论
NaviRAG构建知识分层结构,LLM agent通过迭代识别信息缺口并从适当粒度层级检索相关内容。
原文摘要
Retrieval-augmented generation (RAG) typically relies on a flat retrieval paradigm that maps queries directly to static, isolated text segments. This approach struggles with more complex tasks that require the conditional retrieval and dynamic synthesis of information across different levels of granularity (e.g., from broad concepts to specific evidence). To bridge this gap, we introduce NaviRAG, a novel framework that shifts from passive segment retrieval to active knowledge navigation. NaviRAG first structures the knowledge documents into a hierarchical form, preserving semantic relationships from coarse-grained topics to fine-grained details. Leveraging this reorganized knowledge records, a large language model (LLM) agent actively navigates the records, iteratively identifying information gaps and retrieving relevant content from the most appropriate granularity level. Extensive experiments on long-document QA benchmarks show that NaviRAG consistently improves both retrieval recall and end-to-end answer performance over conventional RAG baselines. Ablation studies confirm performance gains stem from our method's capacity for multi-granular evidence localization and dynamic retrieval planning. We further discuss efficiency, applicable scenario, and future directions of our method, hoping to make RAG systems more intelligent and autonomous.