LLM Reasoning 相关度: 7/10

PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer

David Picard, Nicolas Dufour, Lucas Degeorge, Arijit Ghosh, Davide Allegro, Tom Ravaud, Yohann Perron, Corentin Sautier, Zeynep Sonat Baltaci, Fei Meng, Syrine Kalleli, Marta López-Rauhut, Thibaut Loiseau, Ségolène Albouy, Raphael Baena, Elliot Vincent, Loic Landrieu
arXiv: 2604.06129v1 发布: 2026-04-07 更新: 2026-04-07

AI 摘要

论文提出Polynomial Mixer (PoM),一种线性复杂度的自注意力替代方案,用于高效处理长序列。

主要贡献

  • 提出Polynomial Mixer (PoM) 线性复杂度的token混合机制
  • 证明PoM满足上下文映射属性,保持transformer的通用近似能力
  • 在多个领域验证了PoM的性能和效率

方法论

通过学习到的多项式函数将输入tokens聚合为紧凑的表示,从中检索上下文信息,替代传统自注意力。

原文摘要

This paper introduces the Polynomial Mixer (PoM), a novel token mixing mechanism with linear complexity that serves as a drop-in replacement for self-attention. PoM aggregates input tokens into a compact representation through a learned polynomial function, from which each token retrieves contextual information. We prove that PoM satisfies the contextual mapping property, ensuring that transformers equipped with PoM remain universal sequence-to-sequence approximators. We replace standard self-attention with PoM across five diverse domains: text generation, handwritten text recognition, image generation, 3D modeling, and Earth observation. PoM matches the performance of attention-based models while drastically reducing computational cost when working with long sequences. The code is available at https://github.com/davidpicard/pom.

标签

Attention Linear Complexity Polynomial Mixer Long Sequence Modeling

arXiv 分类

cs.CV cs.AI