AI Agents 相关度: 8/10

AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot

Joydeep Biswas, Sheila Schoepp, Gautham Vasan, Anthony Opipari, Arthur Zhang, Zichao Hu, Sebastian Joseph, Matthew Lease, Junyi Jessy Li, Peter Stone, Kiri L. Wagstaff, Matthew E. Taylor, Odest Chadwicke Jenkins
arXiv: 2604.13940v1 发布: 2026-04-15 更新: 2026-04-15

AI 摘要

AAAI-26进行了大规模AI辅助同行评审实验,结果表明AI评审在技术准确性等方面优于人工评审。

主要贡献

  • 首次大规模部署AI辅助同行评审
  • 证明AI可以生成技术上合理的评审
  • 提出新的基准来评估AI评审质量

方法论

结合前沿模型、工具使用和安全措施,通过多阶段流程为所有论文生成评审。

原文摘要

Scientific peer review faces mounting strain as submission volumes surge, making it increasingly difficult to sustain review quality, consistency, and timeliness. Recent advances in AI have led the community to consider its use in peer review, yet a key unresolved question is whether AI can generate technically sound reviews at real-world conference scale. Here we report the first large-scale field deployment of AI-assisted peer review: every main-track submission at AAAI-26 received one clearly identified AI review from a state-of-the-art system. The system combined frontier models, tool use, and safeguards in a multi-stage process to generate reviews for all 22,977 full-review papers in less than a day. A large-scale survey of AAAI-26 authors and program committee members showed that participants not only found AI reviews useful, but actually preferred them to human reviews on key dimensions such as technical accuracy and research suggestions. We also introduce a novel benchmark and find that our system substantially outperforms a simple LLM-generated review baseline at detecting a variety of scientific weaknesses. Together, these results show that state-of-the-art AI methods can already make meaningful contributions to scientific peer review at conference scale, opening a path toward the next generation of synergistic human-AI teaming for evaluating research.

标签

AI辅助 同行评审 大规模实验

arXiv 分类

cs.AI