LLM Memory & RAG 相关度: 8/10

Short Data, Long Context: Distilling Positional Knowledge in Transformers

Patrick Huber, Ernie Chang, Chinnadhurai Sankar, Rylan Conway, Igor Fedorov, Md Rifat Arefin, Adithya Sagar
arXiv: 2604.06070v1 发布: 2026-04-07 更新: 2026-04-07

AI 摘要

该论文提出一种通过知识蒸馏将长上下文检索能力迁移到短上下文训练模型的方法。

主要贡献

  • 提出基于logits的知识蒸馏可以迁移长上下文能力
  • 分析了RoPE在知识蒸馏中的作用
  • 揭示了长上下文扩展中query状态的结构化更新模式

方法论

通过logits知识蒸馏,在短上下文样本上训练学生模型,并分析RoPE和query状态的更新模式。

原文摘要

Extending the context window of language models typically requires expensive long-context pre-training, posing significant challenges for both training efficiency and data collection. In this paper, we present evidence that long-context retrieval capabilities can be transferred to student models through logit-based knowledge distillation, even when training exclusively on packed short-context samples within a long-context window. We provide comprehensive insights through the lens of Rotary Position Embedding (RoPE) and establish three key findings. First, consistent with prior work, we show that phase-wise RoPE scaling, which maximizes rotational spectrum utilization at each training stage, also achieves the best long-context performance in knowledge distillation setups. Second, we demonstrate that logit-based knowledge distillation can directly enable positional information transfer. Using an experimental setup with packed repeated token sequences, we trace the propagation of positional perturbations from query and key vectors through successive transformer layers to output logits, revealing that positional information systematically influences the teacher's output distribution and, in turn, the distillation signal received by the student model. Third, our analysis uncovers structured update patterns in the query state during long-context extension, with distinct parameter spans exhibiting strong sensitivity to long-context training.

标签

知识蒸馏 长上下文 位置编码 Transformer

arXiv 分类

cs.CL cs.LG