Scaling-Aware Data Selection for End-to-End Autonomous Driving Systems
AI 摘要
提出了MOSAIC框架,通过缩放感知数据选择策略,提升端到端自动驾驶系统性能。
主要贡献
- 提出了MOSAIC数据选择框架
- 将数据划分为域并拟合缩放定律
- 在自动驾驶任务上验证了有效性
方法论
MOSAIC框架通过划分数据域、拟合缩放定律和迭代优化数据混合比例来选择数据。
原文摘要
Large-scale deep learning models for physical AI applications depend on diverse training data collection efforts. These models and correspondingly, the training data, must address different evaluation criteria necessary for the models to be deployable in real-world environments. Data selection policies can guide the development of the training set, but current frameworks do not account for the ambiguity in how data points affect different metrics. In this work, we propose Mixture Optimization via Scaling-Aware Iterative Collection (MOSAIC), a general data selection framework that operates by: (i) partitioning the dataset into domains; (ii) fitting neural scaling laws from each data domain to the evaluation metrics; and (iii) optimizing a data mixture by iteratively adding data from domains that maximize the change in metrics. We apply MOSAIC to autonomous driving (AD), where an End-to-End (E2E) planner model is evaluated on the Extended Predictive Driver Model Score (EPDMS), an aggregate of driving rule compliance metrics. Here, MOSAIC outperforms a diverse set of baselines on EPDMS with up to 80\% less data.