Improving Sparse Autoencoder with Dynamic Attention
AI 摘要
提出了基于动态稀疏注意力机制的稀疏自编码器,提升了解释性和重建效果。
主要贡献
- 提出了基于Cross-Attention和Sparsemax的稀疏自编码器
- 利用Sparsemax动态调整神经元的稀疏程度
- 在重建损失和概念质量上都取得了提升
方法论
使用Cross-Attention结构,latent特征作为query,学习到的字典作为key和value,利用Sparsemax实现动态稀疏激活。
原文摘要
Recently, sparse autoencoders (SAEs) have emerged as a promising technique for interpreting activations in foundation models by disentangling features into a sparse set of concepts. However, identifying the optimal level of sparsity for each neuron remains challenging in practice: excessive sparsity can lead to poor reconstruction, whereas insufficient sparsity may harm interpretability. While existing activation functions such as ReLU and TopK provide certain sparsity guarantees, they typically require additional sparsity regularization or cherry-picked hyperparameters. We show in this paper that dynamically sparse attention mechanisms using sparsemax can bridge this trade-off, due to their ability to determine the activation numbers in a data-dependent manner. Specifically, we first explore a new class of SAEs based on the cross-attention architecture with the latent features as queries and the learnable dictionary as the key and value matrices. To encourage sparse pattern learning, we employ a sparsemax-based attention strategy that automatically infers a sparse set of elements according to the complexity of each neuron, resulting in a more flexible and general activation function. Through comprehensive evaluation and visualization, we show that our approach successfully achieves lower reconstruction loss while producing high-quality concepts, particularly in top-n classification tasks.