LLM Reasoning 相关度: 9/10

Understanding Performance Gap Between Parallel and Sequential Sampling in Large Reasoning Models

Xiangming Gu, Soham De, Larisa Markeeva, Petar Veličković, Razvan Pascanu
arXiv: 2604.05868v1 发布: 2026-04-07 更新: 2026-04-07

AI 摘要

该论文分析了大型推理模型中并行采样优于串行采样的原因,探究了探索不足的影响。

主要贡献

  • 对比了并行采样和串行采样
  • 提出了三种假设来解释性能差距
  • 验证了探索不足是性能差距的主要原因

方法论

通过在不同模型和数据集上进行实验,验证了关于并行采样优于串行采样的假设。

原文摘要

Large Reasoning Models (LRMs) have shown remarkable performance on challenging questions, such as math and coding. However, to obtain a high quality solution, one may need to sample more than once. In principal, there are two sampling strategies that can be composed to form more complex processes: sequential sampling and parallel sampling. In this paper, we first compare these two approaches with rigor, and observe, aligned with previous works, that parallel sampling seems to outperform sequential sampling even though the latter should have more representation power. To understand the underline reasons, we make three hypothesis on the reason behind this behavior: (i) parallel sampling outperforms due to the aggregator operator; (ii) sequential sampling is harmed by needing to use longer contexts; (iii) sequential sampling leads to less exploration due to conditioning on previous answers. The empirical evidence on various model families and sizes (Qwen3, DeepSeek-R1 distilled models, Gemini 2.5) and question domains (math and coding) suggests that the aggregation and context length do not seem to be the main culprit behind the performance gap. In contrast, the lack of exploration seems to play a considerably larger role, and we argue that this is one main cause for the performance gap.

标签

LLM Reasoning Sampling Exploration

arXiv 分类

cs.CL