LongCoT: Benchmarking Long-Horizon Chain-of-Thought Reasoning
AI 摘要
LongCoT基准测试旨在评估LLM在长程链式推理方面的能力,揭示了当前模型在该领域的显著差距。
主要贡献
- 提出了LongCoT基准测试,用于评估长程推理能力
- 设计了跨多个领域(化学、数学、计算机科学等)的2500个专家设计问题
- 揭示了当前最先进模型在长程推理方面的局限性
方法论
构建包含大量推理步骤的问题集,问题需要模型逐步推理,通过验证最终答案评估模型性能。
原文摘要
As language models are increasingly deployed for complex autonomous tasks, their ability to reason accurately over longer horizons becomes critical. An essential component of this ability is planning and managing a long, complex chain-of-thought (CoT). We introduce LongCoT, a scalable benchmark of 2,500 expert-designed problems spanning chemistry, mathematics, computer science, chess, and logic to isolate and directly measure the long-horizon CoT reasoning capabilities of frontier models. Problems consist of a short input with a verifiable answer; solving them requires navigating a graph of interdependent steps that span tens to hundreds of thousands of reasoning tokens. Each local step is individually tractable for frontier models, so failures reflect long-horizon reasoning limitations. At release, the best models achieve <10% accuracy (GPT 5.2: 9.8%; Gemini 3 Pro: 6.1%) on LongCoT, revealing a substantial gap in current capabilities. Overall, LongCoT provides a rigorous measure of long-horizon reasoning, tracking the ability of frontier models to reason reliably over extended periods.