Agent Tuning & Optimization 相关度: 9/10

Synthetic Sandbox for Training Machine Learning Engineering Agents

Yuhang Zhou, Lizhu Zhang, Yifan Wu, Jiayi Liu, Xiangjun Fan, Zhuokai Zhao, Hong Yan
arXiv: 2604.04872v1 发布: 2026-04-06 更新: 2026-04-06

AI 摘要

提出SandMLE框架,通过生成小规模合成数据,加速MLE Agent的on-policy强化学习训练。

主要贡献

  • 提出SandMLE框架,降低了MLE Agent训练的计算成本
  • 实现了MLE Agent的on-policy强化学习
  • 在MLE-bench-lite和MLE-Dojo上验证了有效性

方法论

构建多智能体框架,从小规模种子任务生成多样化的合成MLE环境,将数据集限制在微型规模,进行on-policy强化学习。

原文摘要

As large language model agents advance beyond software engineering (SWE) tasks toward machine learning engineering (MLE), verifying agent behavior becomes orders of magnitude more expensive: while SWE tasks can be verified via fast-executing unit tests, MLE verification requires running full ML pipelines -- data preprocessing, model training, and metric evaluation -- on large datasets at each rollout step, rendering trajectory-wise on-policy reinforcement learning (RL) prohibitively slow. Existing approaches retreat to supervised fine-tuning (SFT) or offline proxy rewards, sacrificing the exploration and generalization benefits of on-policy RL. We observe that sandbox data size is the primary source of this bottleneck. Based on this insight, we introduce SandMLE, a multi-agent framework that generates diverse, verifiable synthetic MLE environments from a small number of seed tasks, preserving the structural and technical complexity of real-world problems while constraining datasets to micro-scale (each task is paired with only 50-200 training samples). Through extensive experiments, we show that SandMLE reduces execution time by over 13 times, enabling large-scale, on-policy trajectory-wise RL for the first time in the MLE domain. On MLE-bench-lite, SandMLE yields significant gains over SFT baselines across Qwen3-8B, 14B, and 30B-A3B, with relative medal rate improvements ranging from 20.3% to 66.9%. Furthermore, the trained policy generalizes across unseen agentic scaffolds, achieving up to 32.4% better HumanRank score on MLE-Dojo.

标签

MLE Agent 强化学习 合成数据 On-policy RL

arXiv 分类

cs.CL cs.LG