UMR: Universal Manipulation Representation

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited cross-embodiment transferability of existing policies caused by their reliance on specific embodiments. We propose the Universal Manipulation Representation (UMR), which decouples actions into embodiment-agnostic world flows and embodiment-specific ego trajectories. By introducing a novel geometric coupling mechanism and SE(3) conjugation techniques, combined with point cloud editing to preserve contact geometry, we construct a compact dual-stream WEPVLA model that enables zero-shot skill transfer from human demonstrations to heterogeneous robots. Experimental results demonstrate that our method achieves success rates of 97.5% and 85.7% in the LIBERO and RLBench simulations, respectively. Furthermore, in real-world scenarios, it attains an average success rate of 91.7% using only a few demonstrations.
📝 Abstract
General-purpose embodied manipulation hinges on a unified action representation that generalizes across embodiments and scales readily. Yet existing policies rely on embodiment-specific action spaces, making cross-embodiment demonstrations difficult to leverage at scale and limiting transfer to new embodiments and spatial variations. To this end, we introduce Universal Manipulation Representation (UMR), a unified action representation that enables zero-shot skill transfer from human demonstrations to heterogeneous robots. UMR decomposes manipulation into two functionally distinct yet geometrically linked components: embodiment-agnostic World Flow, which describes task-relevant object motion in the world frame, and Ego Trajectory, which represents end-effector motion relative to the current pose. We instantiate UMR as World--Ego Point VLA (WEPVLA), a compact 0.5B-parameter policy that learns in the unified geometric action space through a dual-stream Point Action Adapter and a unified Point Action Expert, with an $SE(3)$ conjugation coupling the two components. To improve data efficiency, we complement UMR with a Data-Efficient Strategy (DES) that diversifies object configurations through stage-aware point-cloud editing while preserving demonstrated contact geometry. In simulation, WEPVLA achieves average success rates of 97.5\% on LIBERO and 85.7\% on the 10-task RLBench benchmark. In real-world experiments, a single policy trained on human demonstrations augmented by DES transfers zero-shot to diverse deployment conditions. With about 10 minutes of collected human demonstrations per task and no robot demonstrations, it achieves 91.7\% average success across six evaluation settings, compared with 60.8\% for HumanEgo. Code and additional materials are available at https://umr-wepvla.github.io/.
Problem

Research questions and friction points this paper is trying to address.

Embodied Manipulation
Universal Action Representation
Cross-embodiment Transfer
Zero-shot Skill Transfer
Innovation

Methods, ideas, or system contributions that make the work stand out.

Universal Manipulation Representation
Zero-shot Skill Transfer
Vision-Language-Action Model
Point Cloud Editing
SE(3) Conjugation
💼 Related Jobs
No related jobs found.
S
Song Liu
University of Science and Technology of China, Hefei, China; Suzhou Artificial Intelligence Laboratory, Suzhou, China
L
Linyi Li
University of Science and Technology of China, Hefei, China; Suzhou Artificial Intelligence Laboratory, Suzhou, China
Y
Yanshun Zhao
University of Science and Technology of China, Hefei, China
R
Rxuan Li
University of Science and Technology of China, Hefei, China
X
Xinrui Xu
University of Science and Technology of China, Hefei, China
Yi Ju
Yi Ju
Systems Engineering, UC Berkeley
electric vehiclessustainable infrastructuressmart gridbuilt environment
Y
Yahui Deng
University of Science and Technology of China, Hefei, China
S
Senge Zhang
University of Science and Technology of China, Hefei, China
G
Guoyu Liu
University of Science and Technology of China, Hefei, China
Y
Yixuan Li
University of Science and Technology of China, Hefei, China
Wuyang Zhang
Wuyang Zhang
University of Science and Technology of China
Efficient AI
Yao Li
Yao Li
University of Science and Technology of China
Video CodingVideo Processing
Congcong Zhu
Congcong Zhu
USTC
Multimedia Understanding
J
Jingrun Chen
University of Science and Technology of China, Hefei, China; Suzhou Artificial Intelligence Laboratory, Suzhou, China