LLM Reasoning 相关度: 9/10

Early Stopping for Large Reasoning Models via Confidence Dynamics

Parsa Hosseini, Sumit Nawathe, Mahdi Salmani, Meisam Razaviyayn, Soheil Feizi
arXiv: 2604.04930v1 发布: 2026-04-06 更新: 2026-04-06

AI 摘要

提出CoDE-Stop方法,利用中间答案置信度动态,提前终止大模型推理,提高效率。

主要贡献

  • 提出了CoDE-Stop早期停止方法
  • 分析了推理过程中置信度动态的变化
  • 在多个推理和科学基准测试中验证了方法的有效性

方法论

通过观察正确和错误推理轨迹中置信度的变化,设计算法在置信度达到一定标准时停止推理。

原文摘要

Large reasoning models rely on long chain-of-thought generation to solve complex problems, but extended reasoning often incurs substantial computational cost and can even degrade performance due to overthinking. A key challenge is determining when the model should stop reasoning and produce the final answer. In this work, we study the confidence of intermediate answers during reasoning and observe two characteristic behaviors: correct reasoning trajectories often reach high-confidence answers early, while incorrect rollouts tend to produce long, unproductive reasoning traces and exhibit less reliable confidence dynamics. Motivated by these observations, we propose CoDE-Stop (Confidence Dynamics Early Stop), an early stopping method that leverages the dynamics of intermediate answer confidence to decide when to terminate reasoning, requiring no additional training and easily integrating into existing models. We evaluate CoDE-Stop on diverse reasoning and science benchmarks across multiple models. Compared to prior early stopping methods, it achieves a more favorable accuracy-compute tradeoff and reduces total token usage by 25-50% compared to standard full-length reasoning. In addition, we provide analyses of confidence dynamics during reasoning, offering insights into how confidence changes in both correct and incorrect trajectories.

标签

LLM 推理 Chain-of-Thought 早期停止 置信度

arXiv 分类

cs.CL cs.AI cs.LG