LLM Reasoning 相关度: 9/10

Solving Physics Olympiad via Reinforcement Learning on Physics Simulators

Mihir Prabhudesai, Aryan Satpathy, Yangmin Li, Zheyang Qin, Nikash Bhardwaj, Amir Zadeh, Chuan Li, Katerina Fragkiadaki, Deepak Pathak
arXiv: 2604.11805v1 发布: 2026-04-13 更新: 2026-04-13

AI 摘要

利用物理模拟器生成数据,通过强化学习训练LLM解决物理奥赛问题,实现了sim-to-real迁移。

主要贡献

  • 提出利用物理模拟器生成数据训练LLM的方法
  • 验证了强化学习在物理推理任务中的有效性
  • 证明了sim-to-real迁移在物理问题上的可行性

方法论

使用物理引擎生成合成数据,通过强化学习训练LLM,并在真实物理奥赛题上进行零样本测试。

原文摘要

We have witnessed remarkable advances in LLM reasoning capabilities with the advent of DeepSeek-R1. However, much of this progress has been fueled by the abundance of internet question-answer (QA) pairs, a major bottleneck going forward, since such data is limited in scale and concentrated mainly in domains like mathematics. In contrast, other sciences such as physics lack large-scale QA datasets to effectively train reasoning-capable models. In this work, we show that physics simulators can serve as a powerful alternative source of supervision for training LLMs for physical reasoning. We generate random scenes in physics engines, create synthetic question-answer pairs from simulated interactions, and train LLMs using reinforcement learning on this synthetic data. Our models exhibit zero-shot sim-to-real transfer to real-world physics benchmarks: for example, training solely on synthetic simulated data improves performance on IPhO (International Physics Olympiad) problems by 5-10 percentage points across model sizes. These results demonstrate that physics simulators can act as scalable data generators, enabling LLMs to acquire deep physical reasoning skills beyond the limitations of internet-scale QA data. Code available at: https://sim2reason.github.io/.

标签

Reinforcement Learning Physics Simulation LLM Reasoning

arXiv 分类

cs.LG cs.AI cs.CV cs.RO