LLM Memory & RAG 相关度: 9/10

Transactional Attention: Semantic Sponsorship for KV-Cache Retention

Abhinaba Basu
arXiv: 2604.11288v1 发布: 2026-04-13 更新: 2026-04-13

AI 摘要

Transactional Attention通过结构锚定模式保护关键token免受KV-cache驱逐,显著提高信息检索准确率。

主要贡献

  • 提出Transactional Attention机制,解决KV-cache压缩中的关键信息丢失问题
  • 设计了 attention-free 变体 TA-Fast,降低内存开销
  • 实验证明TA在credential retrieval方面优于现有方法

方法论

利用结构锚定模式(如“key:”)作为sponsor,保护相邻的value-bearing token不被驱逐,提升关键信息的保留。

原文摘要

At K=16 tokens (0.4% of a 4K context), every existing KV-cache compression method achieves 0% on credential retrieval. The failure mode is dormant tokens: credentials, API keys, and configuration values that receive near-zero attention but become essential at generation time. Because these tokens lack the statistical signals that eviction policies rely on, no method based on attention scores, reconstruction loss, or learned retention gates retains them. We introduce Transactional Attention (TA), a sponsorship mechanism in which structural anchor patterns (e.g., "key:", "password:") protect adjacent value-bearing tokens from eviction. TA achieves 100% credential retrieval at K=16 where six baselines (H2O, TOVA, SnapKV, StreamingLLM, PyramidKV, DynamicKV) achieve 0%, and sustains 100% accuracy across 200 function-calling trials. TA-Fast, an attention-free variant, reduces memory overhead by 52% and is compatible with SDPA and FlashAttention. TA is orthogonal to existing compression methods and adds less than 1% latency overhead.

标签

KV-cache Attention机制 内存优化 长文本处理 信息检索

arXiv 分类

cs.CL cs.LG