LLM Reasoning 相关度: 8/10

Calibrated Speculative Decoding: Frequency-Guided Candidate Selection for Efficient Inference

Xuwen Zhou, Fangxin Liu, Chao Wang, Xiao Zheng, Hao Zheng, Min He, Li Jiang, Haibing Guan
arXiv: 2604.13634v1 发布: 2026-04-15 更新: 2026-04-15

AI 摘要

校准推测解码通过频率引导候选选择加速自回归生成,提升LLM推理效率。

主要贡献

  • 提出Calibrated Speculative Decoding (CSD)框架
  • 引入在线纠错记忆模块和语义一致性门控模块
  • 实验证明CSD优于现有方法,实现2.33倍加速

方法论

CSD利用历史拒绝信息纠正错误,并通过概率比而非精确匹配验证候选,提升推测解码的准确性和效率。

原文摘要

Speculative decoding accelerates autoregressive generation by letting draft tokens bypass full verification, but conventional frameworks suffer from frequent false rejections, particularly when draft models produce semantically correct but lexically divergent outputs. In this paper, we present Calibrated Speculative Decoding (CSD), a training-free framework that recovers valid tokens discarded by standard verification. Guided by the principle of "Frequency-Guided Candidate Selection and Probability-Guarded Acceptance," CSD incorporates two lightweight modules: Online Correction Memory, which aggregates historical rejections to propose recurring divergence patterns as rescue candidates, and Semantic Consistency Gating, which verifies candidate admissibility using probability ratios instead of exact token matching. Our evaluation across diverse large language models demonstrates that CSD outperforms existing methods, achieving a peak throughput speedup of 2.33x. CSD preserves model accuracy across all tasks while further boosting performance on complex reasoning datasets. These results establish CSD as a highly effective, lightweight solution for practical LLM deployments.

标签

推测解码 语言模型 模型加速 推理优化

arXiv 分类

cs.CL cs.LG