UMI-3D: Extending Universal Manipulation Interface from Vision-Limited to 3D Spatial Perception

📅 2026-04-15
📈 Citations: 0
Influential: 0
📄 PDF

career value

206K/year
🤖 AI Summary
Existing UMI systems rely on monocular visual SLAM, which suffers from insufficient robustness under occlusions, dynamic scenes, and tracking failures, limiting their applicability in real-world complex environments. This work presents the first integration of a lightweight, low-cost LiDAR into a wrist-worn manipulation interface, establishing a LiDAR-centric SLAM system enhanced by hardware-synchronized multimodal sensing and a unified spatiotemporal calibration framework. The approach achieves high-precision, robust 3D spatial perception and demonstration capture while preserving the original 2D policy architecture. It substantially improves data quality and task generalization, enabling efficient execution of both standard and complex manipulation tasks—including large-scale deformable and articulated object interactions—in real-world settings, significantly outperforming vision-only UMI systems. All software and hardware are open-sourced to advance embodied intelligence research.

Technology Category

Application Category

📝 Abstract
We present UMI-3D, a multimodal extension of the Universal Manipulation Interface (UMI) for robust and scalable data collection in embodied manipulation. While UMI enables portable, wrist-mounted data acquisition, its reliance on monocular visual SLAM makes it vulnerable to occlusions, dynamic scenes, and tracking failures, limiting its applicability in real-world environments. UMI-3D addresses these limitations by introducing a lightweight and low-cost LiDAR sensor tightly integrated into the wrist-mounted interface, enabling LiDAR-centric SLAM with accurate metric-scale pose estimation under challenging conditions. We further develop a hardware-synchronized multimodal sensing pipeline and a unified spatiotemporal calibration framework that aligns visual observations with LiDAR point clouds, producing consistent 3D representations of demonstrations. Despite maintaining the original 2D visuomotor policy formulation, UMI-3D significantly improves the quality and reliability of collected data, which directly translates into enhanced policy performance. Extensive real-world experiments demonstrate that UMI-3D not only achieves high success rates on standard manipulation tasks, but also enables learning of tasks that are challenging or infeasible for the original vision-only UMI setup, including large deformable object manipulation and articulated object operation. The system supports an end-to-end pipeline for data acquisition, alignment, training, and deployment, while preserving the portability and accessibility of the original UMI. All hardware and software components are open-sourced to facilitate large-scale data collection and accelerate research in embodied intelligence: \href{https://umi-3d.github.io}{https://umi-3d.github.io}.
Problem

Research questions and friction points this paper is trying to address.

occlusion
dynamic scenes
tracking failure
monocular visual SLAM
embodied manipulation
Innovation

Methods, ideas, or system contributions that make the work stand out.

LiDAR-centric SLAM
multimodal sensing
spatiotemporal calibration
3D spatial perception
embodied manipulation