RLHND: Video Foundation Models as Physically Grounded Hand Trackers for Robot Learning

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing hand tracking methods that lack physical priors, rendering them inadequate for providing tactile information to support robotic policy learning. Building upon video foundation models, this work proposes a framework to jointly estimate hand pose and veridical tactile information from monocular egocentric videos. Methodologically, it employs the Cosmos 3 diffusion backbone as a deterministic feature extractor, incorporating anatomical constraints and cross-video shape consistency to ensure physical plausibility. Furthermore, linear blend skinning (LBS) feature propagation is introduced to enhance computational efficiency. The proposed approach achieves state-of-the-art performance in both pose and contact force estimation across multiple benchmarks. Its practical utility in robotic manipulation is further validated through motion retargeting and real-world robot experiments.
📝 Abstract
Recently, approaches that leverage human video datasets for robot policy training have become increasingly prevalent. However, most existing hand trackers regress pose from cropped frames with limited priors on hand motion and object interaction, resulting in inaccurate and physically inconsistent estimates. Moreover, the lack of physical cues, e.g., contact and force, limits the use of human videos for robot policy training. To this end, we propose RLHND, a video foundation model-based hand tracking model that jointly estimates hand pose and realistic tactile information from monocular egocentric videos. RLHND turns the pre-trained Cosmos 3 video diffusion backbone into a deterministic clip-level feature extractor via clean-latent conditioning, carrying its learned priors on hand motion and hand-object interaction into tracking. For pose estimation, RLHND (i) predicts hand poses with anatomically plausible joint angles and (ii) enables optional conditioning on the shape parameter to maintain consistent hand shape within the same video and even across videos recorded by the same actor. For tactile estimation, a separate tactile expert stream, trained with the pose stream frozen, predicts dense contact and force over the hand surface. We further adopt LBS-based feature spreading to enable vertex-wise feature extraction without costly per-vertex attention. RLHND achieves state-of-the-art performance across various benchmark datasets for pose estimation, while also achieving state-of-the-art performance in contact and force estimation. Moreover, we demonstrate the utility of RLHND for robot learning through retargeting results and real-world robot experiments. The code will be publicly available at https://seungjun-moon.github.io/rlhnd/.
Problem

Research questions and friction points this paper is trying to address.

hand tracking
robot learning
physical grounding
tactile estimation
pose estimation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Video Foundation Model
Hand Tracking
Tactile Estimation
Clean-Latent Conditioning
Robot Learning