PolicyLLM: Towards Excellent Comprehension of Public Policy for Large Language Models
AI 摘要
提出了PolicyBench基准和PolicyMoE模型,旨在提升LLM对公共政策的理解和应用能力。
主要贡献
- 构建了大规模跨系统的政策理解基准PolicyBench
- 提出了领域专用的混合专家模型PolicyMoE
- 揭示了现有LLM在政策理解方面的局限性
方法论
构建PolicyBench基准,评估LLM在记忆、理解和应用三个认知层次的能力。设计PolicyMoE模型,针对不同认知层次训练专家模块。
原文摘要
Large Language Models (LLMs) are increasingly integrated into real-world decision-making, including in the domain of public policy. Yet, their ability to comprehend and reason about policy-related content remains underexplored. To fill this gap, we present \textbf{\textit{PolicyBench}}, the first large-scale cross-system benchmark (US-China) evaluating policy comprehension, comprising 21K cases across a broad spectrum of policy areas, capturing the diversity and complexity of real-world governance. Following Bloom's taxonomy, the benchmark assesses three core capabilities: (1) \textbf{Memorization}: factual recall of policy knowledge, (2) \textbf{Understanding}: conceptual and contextual reasoning, and (3) \textbf{Application}: problem-solving in real-life policy scenarios. Building on this benchmark, we further propose \textbf{\textit{PolicyMoE}}, a domain-specialized Mixture-of-Experts (MoE) model with expert modules aligned to each cognitive level. The proposed models demonstrate stronger performance on application-oriented policy tasks than on memorization or conceptual understanding, and yields the highest accuracy on structured reasoning tasks. Our results reveal key limitations of current LLMs in policy understanding and suggest paths toward more reliable, policy-focused models.