LLM Reasoning 相关度: 7/10

Ordinary Least Squares is a Special Case of Transformer

Xiaojun Tan, Yuchen Zhao
arXiv: 2604.13656v1 发布: 2026-04-15 更新: 2026-04-15

AI 摘要

论文证明了线性Transformer的特例等价于普通最小二乘法(OLS),揭示了Transformer的统计本质。

主要贡献

  • 证明线性Transformer是OLS的泛化
  • 揭示Transformer内部的快慢记忆机制
  • 建立了现代深度架构和经典统计推断之间的联系

方法论

通过严格的代数证明,利用经验协方差矩阵的谱分解,构造参数设置,使注意力机制等价于OLS。

原文摘要

The statistical essence of the Transformer architecture has long remained elusive: Is it a universal approximator, or a neural network version of known computational algorithms? Through rigorous algebraic proof, we show that the latter better describes Transformer's basic nature: Ordinary Least Squares (OLS) is a special case of the single-layer Linear Transformer. Using the spectral decomposition of the empirical covariance matrix, we construct a specific parameter setting where the attention mechanism's forward pass becomes mathematically equivalent to the OLS closed-form projection. This means attention can solve the problem in one forward pass, not by iterating. Building upon this prototypical case, we further uncover a decoupled slow and fast memory mechanism within Transformers. Finally, the evolution from our established linear prototype to standard Transformers is discussed. This progression facilitates the transition of the Hopfield energy function from linear to exponential memory capacity, thereby establishing a clear continuity between modern deep architectures and classical statistical inference.

标签

Transformer OLS 线性代数 统计推断 神经网络

arXiv 分类

cs.LG cs.AI math.ST stat.ML