LLM Reasoning 相关度: 10/10

LongCoT: Benchmarking Long-Horizon Chain-of-Thought Reasoning

Sumeet Ramesh Motwani, Daniel Nichols, Charles London, Peggy Li, Fabio Pizzati, Acer Blake, Hasan Hammoud, Tavish McDonald, Akshat Naik, Alesia Ivanova, Vignesh Baskaran, Ivan Laptev, Ruben Glatt, Tal Ben-Nun, Philip Torr, Natasha Jaques, Ameya Prabhu, Brian Bartoldson, Bhavya Kailkhura, Christian Schroeder de Witt
arXiv: 2604.14140v1 发布: 2026-04-15 更新: 2026-04-15

AI 摘要

LongCoT基准测试旨在评估LLM在长程链式推理方面的能力,揭示了当前模型在该领域的显著差距。

主要贡献

  • 提出了LongCoT基准测试,用于评估长程推理能力
  • 设计了跨多个领域(化学、数学、计算机科学等)的2500个专家设计问题
  • 揭示了当前最先进模型在长程推理方面的局限性

方法论

构建包含大量推理步骤的问题集,问题需要模型逐步推理,通过验证最终答案评估模型性能。

原文摘要

As language models are increasingly deployed for complex autonomous tasks, their ability to reason accurately over longer horizons becomes critical. An essential component of this ability is planning and managing a long, complex chain-of-thought (CoT). We introduce LongCoT, a scalable benchmark of 2,500 expert-designed problems spanning chemistry, mathematics, computer science, chess, and logic to isolate and directly measure the long-horizon CoT reasoning capabilities of frontier models. Problems consist of a short input with a verifiable answer; solving them requires navigating a graph of interdependent steps that span tens to hundreds of thousands of reasoning tokens. Each local step is individually tractable for frontier models, so failures reflect long-horizon reasoning limitations. At release, the best models achieve <10% accuracy (GPT 5.2: 9.8%; Gemini 3 Pro: 6.1%) on LongCoT, revealing a substantial gap in current capabilities. Overall, LongCoT provides a rigorous measure of long-horizon reasoning, tracking the ability of frontier models to reason reliably over extended periods.

标签

长程推理 链式思考 基准测试 LLM评估

arXiv 分类

cs.LG cs.AI