MARS: Enabling Autoregressive Models Multi-Token Generation
AI 摘要
MARS通过微调AR模型,使其一次预测多个token,提升生成速度和效率。
主要贡献
- 提出MARS微调方法,无需额外参数和架构修改
- 在保证精度不降的前提下,提升生成吞吐量1.5-1.7倍
- 提出块级KV缓存策略,进一步加速推理
方法论
通过在现有指令数据上继续训练,使AR模型能够预测多个token,并结合KV缓存优化推理速度。
原文摘要
Autoregressive (AR) language models generate text one token at a time, even when consecutive tokens are highly predictable given earlier context. We introduce MARS (Mask AutoRegreSsion), a lightweight fine-tuning method that teaches an instruction-tuned AR model to predict multiple tokens per forward pass. MARS adds no architectural modifications, no extra parameters, and produces a single model that can still be called exactly like the original AR model with no performance degradation. Unlike speculative decoding, which maintains a separate draft model alongside the target, or multi-head approaches such as Medusa, which attach additional prediction heads, MARS requires only continued training on existing instruction data. When generating one token per forward pass, MARS matches or exceeds the AR baseline on six standard benchmarks. When allowed to accept multiple tokens per step, it maintains baseline-level accuracy while achieving 1.5-1.7x throughput. We further develop a block-level KV caching strategy for batch inference, achieving up to 1.71x wall-clock speedup over AR with KV cache on Qwen2.5-7B. Finally, MARS supports real-time speed adjustment via confidence thresholding: under high request load, the serving system can increase throughput on the fly without swapping models or restarting, providing a practical latency-quality knob for deployment.