LLM Reasoning 相关度: 7/10

Robust Length Prediction: A Perspective from Heavy-Tailed Prompt-Conditioned Distributions

Jing Wang, Yu-Yang Qian, Ke Xue, Chao Qian, Peng Zhao, Zhi-Hua Zhou
arXiv: 2604.07931v1 发布: 2026-04-09 更新: 2026-04-09

AI 摘要

提出基于重尾分布的LLM输出长度预测方法,解决传统方法将prompt视为唯一目标长度的不足。

主要贡献

  • 揭示prompt-conditioned输出长度分布的重尾特性
  • 提出基于prompt-conditioned长度分布(ProD)的长度预测方法
  • 设计ProD-M和ProD-D两种变体,分别用于稳健的点预测和分布预测

方法论

通过多次生成相同prompt的输出来构建训练目标,利用LLM的隐藏状态,并基于重尾分布进行稳健估计。

原文摘要

Output-length prediction is important for efficient LLM serving, as it directly affects batching, memory reservation, and scheduling. For prompt-only length prediction, most existing methods use a one-shot sampled length as the label, implicitly treating each prompt as if it had one true target length. We show that this is unreliable: even under a fixed model and decoding setup, the same prompt induces a \emph{prompt-conditioned output length distribution}, not a deterministic scalar, and this distribution is consistent with \emph{heavy-tailed} behavior. Motivated by this, we cast length prediction as robust estimation from heavy-tailed prompt-conditioned length distributions. We propose prompt-conditioned length distribution (ProD) methods, which construct training targets from multiple independent generations of the same prompt. Two variants are developed to reuse the served LLM's hidden states: \mbox{ProD-M}, which uses a median-based target for robust point prediction, and ProD-D, which uses a distributional target that preserves prompt-conditioned uncertainty. We provide theoretical justifications by analyzing the estimation error under a surrogate model. Experiments across diverse scenarios show consistent gains in prediction quality.

标签

LLM length prediction heavy-tailed distribution robust estimation

arXiv 分类

cs.LG