🤖 AI Summary
This study addresses the failure of temporal frame alignment when transferring human video pre-trained models to robotic systems due to execution rate discrepancies. To overcome this, we propose the FLTA framework, which incorporates dual priors of global progress and local temporal dynamics. By leveraging ResNet-50 or ViT encoders alongside soft correspondence objectives and temporally constrained optimization, the framework learns a shared cross-modal task progress representation without requiring annotations, effectively handling non-keyframes and varying execution rates. Experimental evaluations demonstrate relative success rate improvements of 46.93%–65.96% in simulation, with real-world manipulation performance significantly surpassing baselines. These results confirm that the proposed approach achieves efficient unsupervised visual adaptation for cross-embodiment transfer.
📝 Abstract
Transferring visual representations pretrained on human videos to robot manipulation requires learning reliable correspondences between human and robot demonstrations. However, paired demonstrations can differ in execution rate and in the proportion of non-key frames that do not directly reflect task progress. Frames at the same relative timestamp may therefore represent different task stages, which can cause correspondence learning to fail. To address these issues, we propose Frame-Level Temporal Alignment (FLTA), a framework that uses two temporal priors to adapt visual encoders pretrained on human videos for robot manipulation. It learns shared task-progress representations without frame-level correspondence annotations, allowing frames at different relative temporal positions to match. A global progress prior combines normalized temporal positions with visual similarity to construct a soft correspondence target. A local temporal order prior penalizes backward transitions while allowing stays and varying forward rates to accommodate execution-rate differences. With ResNet-50 and ViT encoders, our method achieves relative improvements of 46.93% and 65.96%, respectively, over the best baselines in average simulation success rates and achieves higher task success rates on real-world manipulation tasks. These results also suggest that effective human-robot adaptation depends less on the number of parameters updated than on which parameters are selected. Our project page is available at https://rtx5090ultra.github.io/FLTA-Project-Page/.