VidAct: Learning Manipulation from In-the-Wild Videos with Object-Centric 3D Awareness

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenges of limited viewpoints, difficult trajectory transfer, and insufficient 3D geometric awareness in video-based learning from demonstration by proposing an efficient video-to-robot manipulation framework. The framework learns 3D-aware manipulation policies from monocular videos through three key innovations: static object-frame motion canonicalization, residual trajectory transfer, and privileged point cloud supervision. Combined with arbitrary-view mesh reconstruction, it enables adaptation to diverse robot configurations without retargeting. Experiments demonstrate broad applicability across multi-source videos and support zero-shot real-world deployment, significantly enhancing policy generalization and sim-to-real success rates.
📝 Abstract
Video demonstrations offer a scalable alternative to costly robot data for learning manipulation, yet existing reconstruction-based approaches often rely on constrained camera viewpoints or human-to-robot retargeting, while the reconstructed trajectories are difficult to adapt to new objects configurations without distorting the trajectory shape. Another key limitation is that the resulting policies often lack precise object-level 3D geometry awareness, limiting object grounding and object shape awareness critical for precise manipulation. To bridge these gaps, we propose VidAct, an efficient video-to-robot framework that learns object-centric, 3D-aware manipulation policies from a single monocular video per task and enables zero-shot real-world deployment. VidAct consists of three key components. First, VidAct reconstructs object meshes and motion from arbitrary demo videos and canonicalizes the motion in the static object frame, avoiding embodiment-specific retargeting and accommodating diverse camera viewpoints. Second, VidAct employ residual trajectory transfer for adapting the reconstructed motion to novel object configurations while preserving its motion shape. Finally, as the key policy-learning component, VidAct predicts simulation-provided privileged complete-object point clouds at each frame as an auxiliary task while retaining RGB-only deployment, providing dense object-centric supervision over both object pose and 3D geometry. Experiments on human, robot, generated, and internet videos demonstrate broad video applicability and zero-shot deployment. Per-frame complete-object 3D supervision improves policy generalization and sim-to-real success, while residual trajectory transfer enables reliable trajectory adaptation with better shape preservation.
Problem

Research questions and friction points this paper is trying to address.

video-to-robot learning
manipulation policy
3D awareness
trajectory adaptation
object-centric
Innovation

Methods, ideas, or system contributions that make the work stand out.

Object-Centric 3D Awareness
Video-to-Robot Learning
Residual Trajectory Transfer
Zero-Shot Deployment
Complete-Object Point Cloud Supervision
💼 Related Jobs
No related jobs found.