MimicX: Policy-in-the-Loop Supervision Refinement for Video-Driven Humanoid Motion Tracking

πŸ“… 2026-10-06
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the persistent physical execution failures in video-driven humanoid motion tracking, where visually plausible references remain unachievable during real-world deployment. To overcome this challenge, we propose a closed-loop supervised optimization framework that pioneers the transformation of policy execution failures into actionable supervisory signals. By leveraging execution feedback to identify failure modes, the method dynamically adjusts tracking objectives and resets curricula accordingly. Furthermore, it integrates motion retargeting, heterogeneous accelerated execution, and a rollback-verification-based priority selection mechanism to enable iterative policy refinement. Compared with fixed baselines, the proposed approach reduces body tracking error by 25.7% and expands the robust execution horizon by 255.6%, demonstrating substantial improvements in both accuracy and physical feasibility for humanoid motion tracking.
πŸ“ Abstract
Human videos provide rich motion targets for humanoid learning, yet visually plausible references can still produce persistent failures under physics-based execution. These failures reveal where training supervision should change. We present MimicX, a policy-in-the-loop framework that uses execution feedback to refine video-driven humanoid motion tracking. Starting from reconstructed and retargeted motion, MimicX localizes difficult transitions and affected body regions, then jointly adapts tracking objectives and the reset curriculum for policy continuation. Repeated rollout verification selects execution-priority improvements subject to tracking guards. Across four core video tasks, MimicX consistently improves tracking accuracy and Robust Execution Horizon relative to the Fixed Reference baseline. Task-averaged results show a 25.7% reduction in body-tracking error and a 255.6% increase in execution horizon. Additional video, motion-reference, and collision-scene studies evaluate the method beyond the core tasks, while MimicX-HLoop accelerates feedback through heterogeneous execution. Overall, MimicX turns policy failure into actionable supervision for deciding what to refine and which refinement to retain.
Problem

Research questions and friction points this paper is trying to address.

Humanoid motion tracking
Video-driven control
Physics-based execution
Supervision refinement
Innovation

Methods, ideas, or system contributions that make the work stand out.

Policy-in-the-Loop
Humanoid Motion Tracking
Supervision Refinement
Execution Feedback
Reset Curriculum
Shuaijun Liu
Shuaijun Liu
Institute of Software Chinese Academy of Sciences
ε«ζ˜Ÿι€šδΏ‘γ€δΊΊε·₯智能
C
Chenglong Zhang
The Hong Kong University of Science and Technology (Guangzhou)
X
Xuhao Liu
The Hong Kong University of Science and Technology (Guangzhou)
F
Feiyang You
The Hong Kong University of Science and Technology (Guangzhou)
Y
Yifan Liao
The Hong Kong University of Science and Technology (Guangzhou)
S
Shuyang Hao
The Hong Kong University of Science and Technology (Guangzhou)
C
Chaozhe Zhang
The Hong Kong University of Science and Technology (Guangzhou)
C
Chengyu Wu
The Hong Kong University of Science and Technology (Guangzhou)
Zhen Sun
Zhen Sun
DSA Thrust, HKUST(GZ)
LLM security
N
Ningxin Su
The Hong Kong University of Science and Technology (Guangzhou)