Multimodal Learning 相关度: 9/10

UMI-3D: Extending Universal Manipulation Interface from Vision-Limited to 3D Spatial Perception

Ziming Wang
arXiv: 2604.14089v1 发布: 2026-04-15 更新: 2026-04-15

AI 摘要

UMI-3D扩展了UMI,通过集成LiDAR克服了视觉SLAM的局限,提升了数据质量和策略性能。

主要贡献

  • 集成了低成本LiDAR的腕式操作界面
  • 开发了硬件同步的多模态感知流水线
  • 提出了统一的时空标定框架

方法论

通过集成LiDAR,结合视觉信息,进行多模态SLAM,并进行时空标定,提升数据质量和操作策略。

原文摘要

We present UMI-3D, a multimodal extension of the Universal Manipulation Interface (UMI) for robust and scalable data collection in embodied manipulation. While UMI enables portable, wrist-mounted data acquisition, its reliance on monocular visual SLAM makes it vulnerable to occlusions, dynamic scenes, and tracking failures, limiting its applicability in real-world environments. UMI-3D addresses these limitations by introducing a lightweight and low-cost LiDAR sensor tightly integrated into the wrist-mounted interface, enabling LiDAR-centric SLAM with accurate metric-scale pose estimation under challenging conditions. We further develop a hardware-synchronized multimodal sensing pipeline and a unified spatiotemporal calibration framework that aligns visual observations with LiDAR point clouds, producing consistent 3D representations of demonstrations. Despite maintaining the original 2D visuomotor policy formulation, UMI-3D significantly improves the quality and reliability of collected data, which directly translates into enhanced policy performance. Extensive real-world experiments demonstrate that UMI-3D not only achieves high success rates on standard manipulation tasks, but also enables learning of tasks that are challenging or infeasible for the original vision-only UMI setup, including large deformable object manipulation and articulated object operation. The system supports an end-to-end pipeline for data acquisition, alignment, training, and deployment, while preserving the portability and accessibility of the original UMI. All hardware and software components are open-sourced to facilitate large-scale data collection and accelerate research in embodied intelligence: \href{https://umi-3d.github.io}{https://umi-3d.github.io}.

标签

机器人操作 多模态学习 SLAM LiDAR

arXiv 分类

cs.RO cs.AI