LLM Reasoning 相关度: 8/10

Generalization in LLM Problem Solving: The Case of the Shortest Path

Yao Tong, Jiayuan Ye, Anastasia Borovykh, Reza Shokri
arXiv: 2604.15306v1 发布: 2026-04-16 更新: 2026-04-16

AI 摘要

该论文研究了LLM在最短路径问题中的泛化能力,发现空间迁移能力强,但长度缩放泛化失败。

主要贡献

  • 提出了基于最短路径的合成环境,便于控制实验变量
  • 揭示了数据覆盖、强化学习和推理策略对泛化能力的影响
  • 发现了LLM在长度缩放方面存在递归不稳定性问题

方法论

构建合成环境,通过控制训练数据、训练范式和推理策略,评估LLM在空间和长度上的泛化能力。

原文摘要

Whether language models can systematically generalize remains actively debated. Yet empirical performance is jointly shaped by multiple factors such as training data, training paradigms, and inference-time strategies, making failures difficult to interpret. We introduce a controlled synthetic environment based on shortest-path planning, a canonical composable sequential optimization problem. The setup enables clean separation of these factors and supports two orthogonal axes of generalization: spatial transfer to unseen maps and length scaling to longer-horizon problems. We find that models exhibit strong spatial transfer but consistently fail under length scaling due to recursive instability. We further analyze how distinct stages of the learning pipeline influence systematic problem-solving: for example, data coverage sets capability limits; reinforcement learning improves training stability but does not expand those limits; and inference-time scaling enhances performance but cannot rescue length-scaling failures.

标签

LLM Generalization Shortest Path Reinforcement Learning

arXiv 分类

cs.AI cs.LG