3d hand pose estimation

Designs and implements algorithms and systems that estimate the three-dimensional positions and orientations of hand joints and articulated segments from sensor data (e.g., monocular or stereo RGB, depth, or point clouds), producing per-frame or temporally tracked kinematic hand models and joint coordinate outputs. Work includes building neural regressors and model-based optimization/inverse-kinematics pipelines, handling occlusion and self-contact, creating annotation and calibration tools, and defining evaluation metrics for pose accuracy and temporal consistency.

3dhandposeestimation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.54
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

HandOS: 3D Hand Reconstruction in One Stage

Dec 02, 2024
XC
Xingyu Chen
🏛️ Peking University | University of Chinese Academy of Sciences | International Digital Economy Academy

Existing hand reconstruction methods typically adopt a multi-stage paradigm—detection → left/right classification → pose estimation—leading to computational redundancy and error propagation. This paper proposes HandOS, the first end-to-end, single-stage 3D hand reconstruction framework that jointly performs hand detection, 2D keypoint localization, and 3D mesh generation. Its core contributions are: (1) a lightweight end-to-end architecture leveraging a frozen off-the-shelf object detector; (2) a novel interactive 2D–3D decoder that implicitly encodes hand laterality without explicit classification; and (3) a hierarchical attention mechanism jointly optimizing 2D joint locations, 3D mesh vertices, and camera translation parameters. Evaluated on FreiHand, HandOS achieves 5.0 mm PA-MPJPE; on HInt-Ego4D, it attains 64.6% PCK@0.05—setting new state-of-the-art performance.

Eliminates multi-stage hand reconstruction inefficienciesIntegrates 2D and 3D keypoint estimation in one frameworkOvercomes cumulative errors in hand pose estimation

Monocular 3D Hand Pose Estimation with Implicit Camera Alignment

Jun 10, 2025
CP
Christos Pantazopoulos
🏛️ University of Thessaly | Moverse

This paper addresses camera-parameter-free 3D hand pose estimation from a single color image. To tackle challenges including depth ambiguity, occlusion, and anatomical complexity, we propose a calibration-free optimization framework: (1) an implicit camera alignment mechanism that eliminates reliance on intrinsic camera parameters; (2) a fingertip-guided loss to significantly improve distal joint localization accuracy; and (3) an end-to-end differentiable pipeline integrating 2D keypoint supervision, articulated hand priors, and differentiable geometric constraints. Our method achieves state-of-the-art performance on the EgoDexter and Dexter+Object benchmarks and demonstrates strong generalization and robustness on in-the-wild images. The source code is publicly available.

Estimating 3D hand pose from single RGB imagesImproving robustness for in-the-wild hand articulationOvercoming unknown camera parameters in pose estimation

ReJSHand: Efficient Real-Time Hand Pose Estimation and Mesh Reconstruction Using Refined Joint and Skeleton Features

Mar 08, 2025
SA
Shan An
🏛️ Tianjin University | Northeastern University | Beijing University of Technology | Democritus University of Thrace | Tongji University | Southern University of Science and Technology

To address the trade-off between accuracy and efficiency in real-time monocular RGB-based hand pose estimation and mesh reconstruction, this paper proposes a lightweight and efficient framework. Our method introduces a novel joint-skeleton feature refinement mechanism with joint optimization, a feature interaction and expansion module to collaboratively model the 2D-to-3D mapping, and coordinate attention to enhance keypoint representation. It integrates a unified 2D/3D keypoint generator, multi-head self-attention, and linear vertex mapping to improve geometric consistency and inference speed. Evaluated on FreiHand, our approach achieves 72 FPS with PA-MPJPE of 6.3 mm, PA-MPVPE of 6.4 mm, F@05 = 0.756, and F@15 = 0.984—outperforming existing state-of-the-art methods in both accuracy and latency. The framework is particularly suitable for low-latency applications such as robotic dexterous manipulation.

Efficient hand mesh reconstruction from 2D images.High-accuracy hand gesture prediction for responsive interaction.Real-time 3D hand pose estimation for robotics.

This work addresses the significant challenges of 3D hand pose estimation in surgical environments, where strong illumination, occlusions, and uniform glove-wearing lead to highly homogeneous hand appearances, compounded by the scarcity of annotated data. The authors propose the first general-purpose, multi-view 3D hand pose estimation pipeline that requires neither training nor domain-specific fine-tuning. Their approach leverages off-the-shelf pre-trained models to sequentially perform human detection, whole-body pose estimation, and 2D hand keypoint prediction, followed by a multi-view geometric optimization to recover 3D poses. Additionally, they introduce the first large-scale benchmark dataset for 3D hand pose estimation in surgical settings, comprising 68,000 annotated frames. Experiments demonstrate that the proposed method reduces the mean joint error by 31% in 2D and 76% in 3D, substantially outperforming existing baselines.

3D hand pose estimationannotated datasetgloved hands

This work addresses the challenge of accurately estimating the 6D pose of grasped objects under severe occlusion, where vision-only approaches often fail. To overcome this limitation, we propose a multimodal method that fuses visual and fingertip tactile sensing. Tactile signals are uniformly represented as contact point clouds, and a pixel-wise dense visual-tactile feature fusion network is designed to enable high-precision pose estimation. To facilitate training, we extend the NVIDIA DISE-based synthetic data generation pipeline to jointly produce realistic RGB images and corresponding tactile point clouds. Experimental results on a real robotic platform demonstrate that our approach significantly outperforms vision-only baselines, and that the model trained on synthetic data generalizes effectively to real-world scenarios.

6D pose estimationin-hand objectocclusion

Latest Papers

What's happening recently
View more

Existing hand pose datasets are limited in scale, diversity, occlusion handling, arm geometry representation, and RGB-D alignment, hindering model performance and generalization. To address these limitations, this work introduces AnyHand, a large-scale synthetic dataset comprising 2.5 million single-hand and 4.1 million hand–object interaction physically realistic RGB-D images, uniquely providing occlusion annotations, full-arm geometry, and precisely aligned depth data. Furthermore, the authors propose a lightweight, plug-and-play depth fusion module that effectively integrates multimodal features without requiring fine-tuning. Experiments demonstrate that the proposed approach significantly outperforms existing methods on FreiHAND and HO-3D, while exhibiting strong generalization to the unseen domain HO-Cap. Notably, the RGB-D model achieves state-of-the-art performance on HO-3D.

generalizationhand pose estimationocclusion

This work addresses the challenge of maintaining continuous 6D object pose tracking during dexterous manipulation, where visual occlusions caused by fingers often disrupt perception. To overcome this limitation, the authors propose a multimodal approach that fuses proprioception, proximal force/torque, and contact signals with vision. A structured finger-level encoder is introduced to automatically learn a cross-modal gating mechanism, enabling unsupervised disentanglement of translation (vision-dominated) and rotation (tactile-dominated) estimation. By representing tactile inputs as finger-level tokens and employing attention mechanisms, the method effectively integrates sparse tactile cues with visual observations for sequential pose tracking. Experiments demonstrate a 15-fold improvement in pose tracking accuracy under occlusion compared to baseline methods, along with significantly higher success rates in downstream object reorientation tasks, validated on a real robotic system.

6D pose trackinghaptic fusionin-hand manipulation

This work addresses the challenge of robustly reconstructing high-fidelity 4D hand–object interactions under severe occlusion, where existing methods often rely on object templates or physical markers and are sensitive to initial pose estimates. To overcome these limitations, we propose a marker- and template-free multi-view 4D reconstruction framework. Our approach first leverages a multi-view spatio-temporal Transformer to fuse geometric and temporal cues across views, yielding a reliable initialization. Subsequently, we introduce a physics-aware 3D Gaussian optimization mechanism that incorporates tetrahedral deformation constraints, collision detection, and appearance disentanglement to achieve fine-grained reconstruction. Extensive experiments on both public and in-house datasets demonstrate that our method produces robust, artifact-free, high-fidelity results, offering an efficient solution for automated 4D asset generation.

4D hand-object interactionmarkerless captureocclusion

This work addresses the challenge of achieving high-precision closed-loop control in tendon-driven anthropomorphic hands, which typically lack direct joint angle sensing. To circumvent the need for joint encoders, the authors propose a framework for joint angle estimation and control that combines a high-degree-of-freedom kinematic model—formulated using the Denavit–Hartenberg convention—with tendon displacement measurements and a simplified tension model to infer joint angles. A Jacobian-based PI controller augmented with feedforward compensation is then designed for accurate gesture tracking. Validated through nonlinear optimization and MuJoCo simulations on an Anatomically Correct Biomechatronic Hand, the approach successfully reproduces complex gestures with high fidelity, significantly enhancing control performance while preserving mechanical compactness.

anthropomorphic handjoint angle estimationkinematic modeling

Hot Scholars

JD

Jiankang Deng

Imperial College London
Computer VisionMachine Learning
OM

Ozge Mercanoglu Sincan

Center for Vision, Speech and Signal Processing, University of Surrey
Computer VisionDeep Learning
RB

Richard Bowden

Professor of Computer Vision and Machine Learning, CVSSP, University of Surrey
Computer VisionMachine learningArtificial Intelligence
AL

Artem Lykov

PhD student, Skolkovo Institute of Science and Technology
RoboticsAICognitive roboticsVLA