PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer
AI 摘要
论文提出Polynomial Mixer (PoM),一种线性复杂度的自注意力替代方案,用于高效处理长序列。
主要贡献
- 提出Polynomial Mixer (PoM) 线性复杂度的token混合机制
- 证明PoM满足上下文映射属性,保持transformer的通用近似能力
- 在多个领域验证了PoM的性能和效率
方法论
通过学习到的多项式函数将输入tokens聚合为紧凑的表示,从中检索上下文信息,替代传统自注意力。
原文摘要
This paper introduces the Polynomial Mixer (PoM), a novel token mixing mechanism with linear complexity that serves as a drop-in replacement for self-attention. PoM aggregates input tokens into a compact representation through a learned polynomial function, from which each token retrieves contextual information. We prove that PoM satisfies the contextual mapping property, ensuring that transformers equipped with PoM remain universal sequence-to-sequence approximators. We replace standard self-attention with PoM across five diverse domains: text generation, handwritten text recognition, image generation, 3D modeling, and Earth observation. PoM matches the performance of attention-based models while drastically reducing computational cost when working with long sequences. The code is available at https://github.com/davidpicard/pom.