Multimodal Learning 相关度: 9/10

Grounded World Model for Semantically Generalizable Planning

Quanyi Li, Lan Feng, Haonan Zhang, Wuyang Li, Letian Wang, Alexandre Alahi, Harold Soh
arXiv: 2604.11751v1 发布: 2026-04-13 更新: 2026-04-13

AI 摘要

提出了一个基于视觉-语言对齐潜在空间的Grounded World Model,用于语义泛化的规划。

主要贡献

  • 提出Grounded World Model (GWM)
  • 将视觉运动MPC转换为视觉-语言对齐(VLA)问题
  • WISER benchmark评估GWM-MPC的性能

方法论

学习视觉-语言对齐的潜在空间,基于动作结果与任务指令的相似度进行动作评分,实现语义泛化。

原文摘要

In Model Predictive Control (MPC), world models predict the future outcomes of various action proposals, which are then scored to guide the selection of the optimal action. For visuomotor MPC, the score function is a distance metric between a predicted image and a goal image, measured in the latent space of a pretrained vision encoder like DINO and JEPA. However, it is challenging to obtain the goal image in advance of the task execution, particularly in new environments. Additionally, conveying the goal through an image offers limited interactivity compared with natural language. In this work, we propose to learn a Grounded World Model (GWM) in a vision-language-aligned latent space. As a result, each proposed action is scored based on how close its future outcome is to the task instruction, reflected by the similarity of embeddings. This approach transforms the visuomotor MPC to a VLA that surpasses VLM-based VLAs in semantic generalization. On the proposed WISER benchmark, GWM-MPC achieves a 87% success rate on the test set comprising 288 tasks that feature unseen visual signals and referring expressions, yet remain solvable with motions demonstrated during training. In contrast, traditional VLAs achieve an average success rate of 22%, even though they overfit the training set with a 90% success rate.

标签

World Model MPC Vision-Language Alignment Semantic Generalization

arXiv 分类

cs.RO cs.AI