Latent Planning Emerges with Scale
AI 摘要
研究表明,LLM的潜在规划能力随模型规模增大而增强,但规划深度有限。
主要贡献
- 定义了潜在规划并提出衡量框架
- 揭示了LLM中潜在规划机制的存在
- 发现模型规模与潜在规划能力的正相关性
方法论
通过设计简单规划任务(如补全押韵对句),分析Qwen-3系列模型的生成过程,研究内部表征与输出token间的关系。
原文摘要
LLMs can perform seemingly planning-intensive tasks, like writing coherent stories or functioning code, without explicitly verbalizing a plan; however, the extent to which they implicitly plan is unknown. In this paper, we define latent planning as occurring when LLMs possess internal planning representations that (1) cause the generation of a specific future token or concept, and (2) shape preceding context to license said future token or concept. We study the Qwen-3 family (0.6B-14B) on simple planning tasks, finding that latent planning ability increases with scale. Models that plan possess features that represent a planned-for word like "accountant", and cause them to output "an" rather than "a"; moreover, even the less-successful Qwen-3 4B-8B have nascent planning mechanisms. On the more complex task of completing rhyming couplets, we find that models often identify a rhyme ahead of time, but even large models seldom plan far ahead. However, we can elicit some planning that increases with scale when steering models towards planned words in prose. In sum, we offer a framework for measuring planning and mechanistic evidence of how models' planning abilities grow with scale.