StarVLA-$α$: Reducing Complexity in Vision-Language-Action Systems
AI 摘要
StarVLA-α旨在简化VLA系统,通过最小化复杂性实现强大的机器人控制性能,并作为VLA研究的基准。
主要贡献
- 提出了一个简单且强大的VLA基线模型StarVLA-α。
- 系统地研究了动作建模策略、机器人预训练和接口工程等关键设计轴。
- 在多个基准测试中,StarVLA-α的性能具有竞争力,证明了VLM骨干网络的重要性。
方法论
通过简化架构和流程,StarVLA-α专注于VLM骨干网络,并评估不同设计选择对性能的影响,最终实现强大的机器人控制。
原文摘要
Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for building general-purpose robotic agents. However, the VLA landscape remains highly fragmented and complex: as existing approaches vary substantially in architectures, training data, embodiment configurations, and benchmark-specific engineering. In this work, we introduce StarVLA-$α$, a simple yet strong baseline designed to study VLA design choices under controlled conditions. StarVLA-$α$ deliberately minimizes architectural and pipeline complexity to reduce experimental confounders and enable systematic analysis. Specifically, we re-evaluate several key design axes, including action modeling strategies, robot-specific pretraining, and interface engineering. Across unified multi-benchmark training on LIBERO, SimplerEnv, RoboTwin, and RoboCasa, the same simple baseline remains highly competitive, indicating that a strong VLM backbone combined with minimal design is already sufficient to achieve strong performance without relying on additional architectural complexity or engineering tricks. Notably, our single generalist model outperforms $π_{0.5}$ by 20\% on the public real-world RoboChallenge benchmark. We expect StarVLA-$α$ to serve as a solid starting point for future research in the VLA regime. Code will be released at https://github.com/starVLA/starVLA.