Multimodal Learning 相关度: 8/10

LAMP: Lift Image-Editing as General 3D Priors for Open-world Manipulation

Jingjing Wang, Zhengdong Hong, Chong Bao, Yuke Zhu, Junhan Sun, Guofeng Zhang
arXiv: 2604.08475v1 发布: 2026-04-09 更新: 2026-04-09

AI 摘要

LAMP通过提升图像编辑为3D先验,实现开放世界操作中的零样本泛化。

主要贡献

  • 提出LAMP,利用图像编辑的3D先验进行操作。
  • 将图像编辑中的2D空间信息提升为3D变换。
  • 在开放世界操作中实现精确的3D变换和零样本泛化。

方法论

LAMP提取图像编辑中蕴含的3D变换信息,作为连续、几何感知的表示,指导机器人操作。

原文摘要

Human-like generalization in open-world remains a fundamental challenge for robotic manipulation. Existing learning-based methods, including reinforcement learning, imitation learning, and vision-language-action-models (VLAs), often struggle with novel tasks and unseen environments. Another promising direction is to explore generalizable representations that capture fine-grained spatial and geometric relations for open-world manipulation. While large-language-model (LLMs) and vision-language-model (VLMs) provide strong semantic reasoning based on language or annotated 2D representations, their limited 3D awareness restricts their applicability to fine-grained manipulation. To address this, we propose LAMP, which lifts image-editing as 3D priors to extract inter-object 3D transformations as continuous, geometry-aware representations. Our key insight is that image-editing inherently encodes rich 2D spatial cues, and lifting these implicit cues into 3D transformations provides fine-grained and accurate guidance for open-world manipulation. Extensive experiments demonstrate that \codename delivers precise 3D transformations and achieves strong zero-shot generalization in open-world manipulation. Project page: https://zju3dv.github.io/LAMP/.

标签

机器人操作 3D先验 图像编辑 零样本学习 视觉语言模型

arXiv 分类

cs.CV