LLM Reasoning 相关度: 9/10

A Mechanistic Analysis of Looped Reasoning Language Models

Hugh Blayney, Álvaro Arroyo, Johan Obando-Ceron, Pablo Samuel Castro, Aaron Courville, Michael M. Bronstein, Xiaowen Dong
arXiv: 2604.11791v1 发布: 2026-04-13 更新: 2026-04-13

AI 摘要

分析循环推理语言模型,揭示其内部机制和推理阶段与前馈模型的相似性。

主要贡献

  • 发现循环模型每一层收敛到不同固定点,形成循环轨迹
  • 证明循环模型的注意力头行为趋于稳定
  • 揭示循环块学习的推理阶段与前馈模型相似

方法论

对循环语言模型的潜在状态进行机制分析,研究循环递归,并分析注意力头行为。

原文摘要

Reasoning has become a central capability in large language models. Recent research has shown that reasoning performance can be improved by looping an LLM's layers in the latent dimension, resulting in looped reasoning language models. Despite promising results, few works have investigated how their internal dynamics differ from those of standard feedforward models. In this paper, we conduct a mechanistic analysis of the latent states in looped language models, focusing in particular on how the stages of inference observed in feedforward models compare to those observed in looped ones. To this end, we analyze cyclic recurrence and show that for many of the studied models each layer in the cycle converges to a distinct fixed point; consequently, the recurrent block follows a consistent cyclic trajectory in the latent space. We provide evidence that as these fixed points are reached, attention-head behavior stabilizes, leading to constant behavior across recurrences. Empirically, we discover that recurrent blocks learn stages of inference that closely mirror those of feedforward models, repeating these stages in depth with each iteration. We study how recurrent block size, input injection, and normalization influence the emergence and stability of these cyclic fixed points. We believe these findings help translate mechanistic insights into practical guidance for architectural design.

标签

循环推理 语言模型 机制分析 注意力机制

arXiv 分类

cs.LG cs.AI