LLM Reasoning 相关度: 9/10

Reasoning Models Know What's Important, and Encode It in Their Activations

Yaniv Nikankin, Martin Tutek, Tomer Ashuach, Jonathan Rosenfeld, Yonatan Belinkov
arXiv: 2604.18307v1 发布: 2026-04-20 更新: 2026-04-20

AI 摘要

研究表明,语言模型激活中包含比推理链token更多的重要步骤信息,模型内部编码步骤重要性。

主要贡献

  • 证明了模型激活包含比token更多的推理步骤重要性信息
  • 揭示了模型内部存在步骤重要性的表示
  • 该内部表示具有跨模型泛化能力

方法论

通过训练探测器预测激活的重要性,并分析激活与表面特征的相关性,来研究模型内部如何表示步骤重要性。

原文摘要

Language models often solve complex tasks by generating long reasoning chains, consisting of many steps with varying importance. While some steps are crucial for generating the final answer, others are removable. Determining which steps matter most, and why, remains an open question central to understanding how models process reasoning. We investigate if this question is best approached through model internals or through tokens of the reasoning chain itself. We find that model activations contain more information than tokens for identifying important reasoning steps. Crucially, by training probes on model activations to predict importance, we show that models encode an internal representation of step importance, even prior to the generation of subsequent steps. This internal representation of importance generalizes across models, is distributed across layers, and does not correlate with surface-level features, such as a step's relative position or its length. Our findings suggest that analyzing activations can reveal aspects of reasoning that surface-level approaches fundamentally miss, indicating that reasoning analyses should look into model internals.

标签

LLM Reasoning Model Internals

arXiv 分类

cs.CL