Multimodal Learning 相关度: 9/10

StarVLA-$α$: Reducing Complexity in Vision-Language-Action Systems

Jinhui Ye, Ning Gao, Senqiao Yang, Jinliang Zheng, Zixuan Wang, Yuxin Chen, Pengguang Chen, Yilun Chen, Shu Liu, Jiaya Jia
arXiv: 2604.11757v1 发布: 2026-04-13 更新: 2026-04-13

AI 摘要

StarVLA-α旨在简化VLA系统,通过最小化复杂性实现强大的机器人控制性能,并作为VLA研究的基准。

主要贡献

  • 提出了一个简单且强大的VLA基线模型StarVLA-α。
  • 系统地研究了动作建模策略、机器人预训练和接口工程等关键设计轴。
  • 在多个基准测试中,StarVLA-α的性能具有竞争力,证明了VLM骨干网络的重要性。

方法论

通过简化架构和流程,StarVLA-α专注于VLM骨干网络,并评估不同设计选择对性能的影响,最终实现强大的机器人控制。

原文摘要

Vision-Language-Action (VLA) models have recently emerged as a promising paradigm for building general-purpose robotic agents. However, the VLA landscape remains highly fragmented and complex: as existing approaches vary substantially in architectures, training data, embodiment configurations, and benchmark-specific engineering. In this work, we introduce StarVLA-$α$, a simple yet strong baseline designed to study VLA design choices under controlled conditions. StarVLA-$α$ deliberately minimizes architectural and pipeline complexity to reduce experimental confounders and enable systematic analysis. Specifically, we re-evaluate several key design axes, including action modeling strategies, robot-specific pretraining, and interface engineering. Across unified multi-benchmark training on LIBERO, SimplerEnv, RoboTwin, and RoboCasa, the same simple baseline remains highly competitive, indicating that a strong VLM backbone combined with minimal design is already sufficient to achieve strong performance without relying on additional architectural complexity or engineering tricks. Notably, our single generalist model outperforms $π_{0.5}$ by 20\% on the public real-world RoboChallenge benchmark. We expect StarVLA-$α$ to serve as a solid starting point for future research in the VLA regime. Code will be released at https://github.com/starVLA/starVLA.

标签

Vision-Language-Action Robotics Multimodal Learning Reinforcement Learning

arXiv 分类

cs.RO cs.AI cs.CV