HAMSA: Scanning-Free Vision State Space Models via SpectralPulseNet
AI 摘要
HAMSA是一种无扫描的视觉状态空间模型,在频谱域直接操作,提升了速度和效率。
主要贡献
- 提出了一种简化的核参数化方法,使用高斯初始化的复数核代替传统矩阵
- 引入了SpectralPulseNet (SPN)机制,实现输入相关的频率门控
- 设计了Spectral Adaptive Gating Unit (SAGU),用于频率域的稳定梯度流动
方法论
HAMSA利用FFT卷积消除序列扫描,在频谱域直接进行计算,实现了O(L log L)的复杂度。
原文摘要
Vision State Space Models (SSMs) like Vim, VMamba, and SiMBA rely on complex scanning strategies to adapt sequential SSMs to process 2D images, introducing computational overhead and architectural complexity. We propose HAMSA, a scanning-free SSM operating directly in the spectral domain. HAMSA introduces three key innovations: (1) simplified kernel parameterization-a single Gaussian-initialized complex kernel replacing traditional (A, B, C) matrices, eliminating discretization instabilities; (2) SpectralPulseNet (SPN)-an input-dependent frequency gating mechanism enabling adaptive spectral modulation; and (3) Spectral Adaptive Gating Unit (SAGU)-magnitude-based gating for stable gradient flow in the frequency domain. By leveraging FFT-based convolution, HAMSA eliminates sequential scanning while achieving O(L log L) complexity with superior simplicity and efficiency. On ImageNet-1K, HAMSA reaches 85.7% top-1 accuracy (state-of-the-art among SSMs), with 2.2 X faster inference than transformers (4.2ms vs 9.2ms for DeiT-S) and 1.4-1.9X speedup over scanning-based SSMs, while using less memory (2.1GB vs 3.2-4.5GB) and energy (12.5J vs 18-25J). HAMSA demonstrates strong generalization across transfer learning and dense prediction tasks.