Video-Conditioned Generative Joint 2D-3D Hand Motion Recovery

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of low reconstruction accuracy and temporal inconsistency in 3D hand pose estimation caused by video occlusions. To overcome these limitations, this work proposes JoHan, a unified generative framework that departs from conventional frame-by-frame prediction paradigms. Instead, it directly generates spatiotemporally aligned 2D and 3D local poses from video sequences in an end-to-end manner. Specifically, the method leverages 2D spatiotemporal cues to guide 3D reconstruction, learns motion priors to enhance temporal consistency, and recovers global poses through 2D-3D correspondences. By integrating generative modeling with cross-representation learning, JoHan significantly improves both reconstruction accuracy and inference efficiency, achieving high per-frame precision alongside temporally smooth hand motion capture.
📝 Abstract
Recovering faithful 3D hand motion from video remains challenging due to frequent occlusions and incomplete visual observations, which make frame-wise pose estimates unreliable and temporally inconsistent. To address this problem, we propose JoHan, a unified generative framework that recovers hand motion directly from video sequences without relying on intermediate per-frame pose predictions. Trained from scratch, our model jointly generates aligned 2D and 3D local hand pose sequences by learning their temporal dynamics and cross-representation correspondence. The generated 2D trajectories exploit direct spatial and temporal cues from the 2D images to guide the following generative 3D motion reconstruction, while the learned motion prior promotes temporal consistency. Their learned 2D-3D correspondence further enables recovery of the hand's global position and orientation relative to the camera. Extensive experiments on challenging benchmarks demonstrate significantly improved accuracy and speed in local hand-pose and camera-space reconstruction. Notably, our method captures much better hand-motion dynamics, producing significantly smoother motion than previous methods while maintaining high per-frame pose accuracy.
Problem

Research questions and friction points this paper is trying to address.

3D hand motion recovery
occlusion
temporal inconsistency
video-based reconstruction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Generative Framework
Joint 2D-3D Motion Recovery
Video-Conditioned Generation
Temporal Consistency
Cross-Representation Correspondence
🔎 Similar Papers
No similar papers found.