Enabling a Unified Cross-Domain Representation for Two-Finger Gripper Manipulation via Interaction-Centric Modeling

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of cross-embodiment generalization in imitation learning, which arises from the entanglement of task semantics with hardware-specific visual geometry. To overcome this, we propose an interaction-centric unified framework that generates canonical representations through a parameterized universal gripper abstraction. This representation is integrated with Vision-Language Model (VLM) reasoning, SAM 2.1 segmentation, artificial potential fields, and a Flow-Matching Transformer for hybrid feature-based action prediction. The proposed method uniquely unifies competitive benchmark performance with zero-shot Sim-to-Real transfer under extreme cross-embodiment and cross-view conditions. Extensive evaluations on both simulated environments and real-world heterogeneous robotic platforms validate the framework’s robust cross-domain manipulation capabilities.
📝 Abstract
Achieving robust cross-embodiment generalization in imitation learning demands overcoming a critical representation flaw that inextricably entangles task semantics with hardware-specific visual geometry. We propose an interaction-centric framework that leverages the shared structure of two-finger grippers via a parameterized universal gripper abstraction, yielding a canonical gripper-frame representation. Given language and RGB-D observations, a VLM infers the subtask and grounds an interaction triplet (gripper, held, target), while SAM~2.1 tracks masks to reduce VLM queries. We design concise hybrid features that combine target/collision artificial potential fields for global guidance with segmented gripper-frame point clouds for local geometry, and use a Flow-Matching Transformer to predict smooth 7-DoF action chunks. Experiments in simulation and real-world tasks demonstrate that ours is the first imitation learning approach to simultaneously achieve competitive benchmark scores and extreme cross-embodiment/cross-viewpoint zero-shot sim-to-real transfer to completely distinct, heterogeneous robot platforms.
Problem

Research questions and friction points this paper is trying to address.

cross-embodiment generalization
imitation learning
representation flaw
zero-shot transfer
two-finger gripper
Innovation

Methods, ideas, or system contributions that make the work stand out.

Interaction-Centric Modeling
Cross-Embodiment Generalization
Flow-Matching Transformer
Universal Gripper Abstraction
Sim-to-Real Transfer
💼 Related Jobs
No related jobs found.
G
Guanlin Li
Key Laboratory of Data Engineering and Knowledge Engineering, MOE, and School of Information, Renmin University of China, China
S
Shifeng Bao
Key Laboratory of Data Engineering and Knowledge Engineering, MOE, and School of Information, Renmin University of China, China
Y
Yihan Zhao
Key Laboratory of Data Engineering and Knowledge Engineering, MOE, and School of Information, Renmin University of China, China
H
Haitao Shen
Key Laboratory of Data Engineering and Knowledge Engineering, MOE, and School of Information, Renmin University of China, China
H
Haoyang Li
Key Laboratory of Data Engineering and Knowledge Engineering, MOE, and School of Information, Renmin University of China, China
C
Chen Zhao
Key Laboratory of Data Engineering and Knowledge Engineering, MOE, and School of Information, Renmin University of China, China
Tong Yang
Tong Yang
Peking University, Beijing, China. PKU. 北京大学
SketchNetwork measurementBloom filterIP lookupHash Table
Jie Tang
Jie Tang
UW Madison
Computed Tomography
Jing Zhang
Jing Zhang
Renmin University of China
large model alignmentmodel compression & inference optimizationdata intelligence