TriAttention: Efficient Long Reasoning with Trigonometric KV Compression
AI 摘要
TriAttention通过三角函数KV压缩,提升LLM长文本推理效率,降低内存占用。
主要贡献
- 提出TriAttention机制,利用Q/K向量集中特性压缩KV缓存
- 在AIME25数据集上验证了TriAttention的有效性,性能优于现有方法
- 实现了在单GPU上部署OpenClaw等长文本应用
方法论
利用预RoPE空间的Q/K向量集中特性,通过三角函数计算距离偏好,评估Key的重要性,并结合Q/K范数进行优化。
原文摘要
Extended reasoning in large language models (LLMs) creates severe KV cache memory bottlenecks. Leading KV cache compression methods estimate KV importance using attention scores from recent post-RoPE queries. However, queries rotate with position during RoPE, making representative queries very few, leading to poor top-key selection and unstable reasoning. To avoid this issue, we turn to the pre-RoPE space, where we observe that Q and K vectors are highly concentrated around fixed non-zero centers and remain stable across positions -- Q/K concentration. We show that this concentration causes queries to preferentially attend to keys at specific distances (e.g., nearest keys), with the centers determining which distances are preferred via a trigonometric series. Based on this, we propose TriAttention to estimate key importance by leveraging these centers. Via the trigonometric series, we use the distance preference characterized by these centers to score keys according to their positions, and also leverage Q/K norms as an additional signal for importance estimation. On AIME25 with 32K-token generation, TriAttention matches Full Attention reasoning accuracy while achieving 2.5x higher throughput or 10.7x KV memory reduction, whereas leading baselines achieve only about half the accuracy at the same efficiency. TriAttention enables OpenClaw deployment on a single consumer GPU, where long context would otherwise cause out-of-memory with Full Attention.