LLM Reasoning 相关度: 8/10

What Is the Minimum Architecture for Prolepsis? Early Irrevocable Commitment Across Tasks in Small Transformers

Éric Jacopin
arXiv: 2604.15010v1 发布: 2026-04-16 更新: 2026-04-16

AI 摘要

研究Transformer如何及早做出决策并防止修正,提出了“前摄性(prolepsis)”概念。

主要贡献

  • 定义了prolepsis,即Transformer的早期决策和持续承诺机制
  • 发现了规划位点的几何一致性
  • 揭示了特定注意力头在决策路由中的作用

方法论

通过在Gemma和Llama等开放模型上进行实验,分析Transformer的内部机制,并使用多种方法进行验证。

原文摘要

When do transformers commit to a decision, and what prevents them from correcting it? We introduce \textbf{prolepsis}: a transformer commits early, task-specific attention heads sustain the commitment, and no layer corrects it. Replicating \citeauthor{lindsey2025biology}'s (\citeyear{lindsey2025biology}) planning-site finding on open models (Gemma~2 2B, Llama~3.2 1B), we ask five questions. (Q1)~Planning is invisible to six residual-stream methods; CLTs are necessary. (Q2)~The planning-site spike replicates with identical geometry. (Q3)~Specific attention heads route the decision to the output, filling a gap flagged as invisible to attribution graphs. (Q4)~Search requires ${\leq}16$ layers; commitment requires more. (Q5)~Factual recall shows the same motif at a different network depth, with zero overlap between recurring planning heads and the factual top-10. Prolepsis is architectural: the template is shared, the routing substrates differ. All experiments run on a single consumer GPU (16\,GB VRAM).

标签

Transformer prolepsis attention mechanism planning irreversible commitment

arXiv 分类

cs.LG cs.AI cs.CL