Score
Designs and implements algorithms and systems that estimate the three-dimensional positions and orientations of hand joints and articulated segments from sensor data (e.g., monocular or stereo RGB, depth, or point clouds), producing per-frame or temporally tracked kinematic hand models and joint coordinate outputs. Work includes building neural regressors and model-based optimization/inverse-kinematics pipelines, handling occlusion and self-contact, creating annotation and calibration tools, and defining evaluation metrics for pose accuracy and temporal consistency.
Existing hand reconstruction methods typically adopt a multi-stage paradigm—detection → left/right classification → pose estimation—leading to computational redundancy and error propagation. This paper proposes HandOS, the first end-to-end, single-stage 3D hand reconstruction framework that jointly performs hand detection, 2D keypoint localization, and 3D mesh generation. Its core contributions are: (1) a lightweight end-to-end architecture leveraging a frozen off-the-shelf object detector; (2) a novel interactive 2D–3D decoder that implicitly encodes hand laterality without explicit classification; and (3) a hierarchical attention mechanism jointly optimizing 2D joint locations, 3D mesh vertices, and camera translation parameters. Evaluated on FreiHand, HandOS achieves 5.0 mm PA-MPJPE; on HInt-Ego4D, it attains 64.6% PCK@0.05—setting new state-of-the-art performance.
This paper addresses camera-parameter-free 3D hand pose estimation from a single color image. To tackle challenges including depth ambiguity, occlusion, and anatomical complexity, we propose a calibration-free optimization framework: (1) an implicit camera alignment mechanism that eliminates reliance on intrinsic camera parameters; (2) a fingertip-guided loss to significantly improve distal joint localization accuracy; and (3) an end-to-end differentiable pipeline integrating 2D keypoint supervision, articulated hand priors, and differentiable geometric constraints. Our method achieves state-of-the-art performance on the EgoDexter and Dexter+Object benchmarks and demonstrates strong generalization and robustness on in-the-wild images. The source code is publicly available.
To address the trade-off between accuracy and efficiency in real-time monocular RGB-based hand pose estimation and mesh reconstruction, this paper proposes a lightweight and efficient framework. Our method introduces a novel joint-skeleton feature refinement mechanism with joint optimization, a feature interaction and expansion module to collaboratively model the 2D-to-3D mapping, and coordinate attention to enhance keypoint representation. It integrates a unified 2D/3D keypoint generator, multi-head self-attention, and linear vertex mapping to improve geometric consistency and inference speed. Evaluated on FreiHand, our approach achieves 72 FPS with PA-MPJPE of 6.3 mm, PA-MPVPE of 6.4 mm, F@05 = 0.756, and F@15 = 0.984—outperforming existing state-of-the-art methods in both accuracy and latency. The framework is particularly suitable for low-latency applications such as robotic dexterous manipulation.
This work addresses the significant challenges of 3D hand pose estimation in surgical environments, where strong illumination, occlusions, and uniform glove-wearing lead to highly homogeneous hand appearances, compounded by the scarcity of annotated data. The authors propose the first general-purpose, multi-view 3D hand pose estimation pipeline that requires neither training nor domain-specific fine-tuning. Their approach leverages off-the-shelf pre-trained models to sequentially perform human detection, whole-body pose estimation, and 2D hand keypoint prediction, followed by a multi-view geometric optimization to recover 3D poses. Additionally, they introduce the first large-scale benchmark dataset for 3D hand pose estimation in surgical settings, comprising 68,000 annotated frames. Experiments demonstrate that the proposed method reduces the mean joint error by 31% in 2D and 76% in 3D, substantially outperforming existing baselines.
This work addresses the challenge of accurately estimating the 6D pose of grasped objects under severe occlusion, where vision-only approaches often fail. To overcome this limitation, we propose a multimodal method that fuses visual and fingertip tactile sensing. Tactile signals are uniformly represented as contact point clouds, and a pixel-wise dense visual-tactile feature fusion network is designed to enable high-precision pose estimation. To facilitate training, we extend the NVIDIA DISE-based synthetic data generation pipeline to jointly produce realistic RGB images and corresponding tactile point clouds. Experimental results on a real robotic platform demonstrate that our approach significantly outperforms vision-only baselines, and that the model trained on synthetic data generalizes effectively to real-world scenarios.
Existing hand pose datasets are limited in scale, diversity, occlusion handling, arm geometry representation, and RGB-D alignment, hindering model performance and generalization. To address these limitations, this work introduces AnyHand, a large-scale synthetic dataset comprising 2.5 million single-hand and 4.1 million hand–object interaction physically realistic RGB-D images, uniquely providing occlusion annotations, full-arm geometry, and precisely aligned depth data. Furthermore, the authors propose a lightweight, plug-and-play depth fusion module that effectively integrates multimodal features without requiring fine-tuning. Experiments demonstrate that the proposed approach significantly outperforms existing methods on FreiHAND and HO-3D, while exhibiting strong generalization to the unseen domain HO-Cap. Notably, the RGB-D model achieves state-of-the-art performance on HO-3D.
为解决单深度图像3D手姿态估计中的拓扑依赖性和特征空间干扰问题,提出KAD-Net,通过手指拓扑约束模块和任务解耦框架提高估计准确性。
This work addresses the challenge of maintaining continuous 6D object pose tracking during dexterous manipulation, where visual occlusions caused by fingers often disrupt perception. To overcome this limitation, the authors propose a multimodal approach that fuses proprioception, proximal force/torque, and contact signals with vision. A structured finger-level encoder is introduced to automatically learn a cross-modal gating mechanism, enabling unsupervised disentanglement of translation (vision-dominated) and rotation (tactile-dominated) estimation. By representing tactile inputs as finger-level tokens and employing attention mechanisms, the method effectively integrates sparse tactile cues with visual observations for sequential pose tracking. Experiments demonstrate a 15-fold improvement in pose tracking accuracy under occlusion compared to baseline methods, along with significantly higher success rates in downstream object reorientation tasks, validated on a real robotic system.
This work addresses the challenge of robustly reconstructing high-fidelity 4D hand–object interactions under severe occlusion, where existing methods often rely on object templates or physical markers and are sensitive to initial pose estimates. To overcome these limitations, we propose a marker- and template-free multi-view 4D reconstruction framework. Our approach first leverages a multi-view spatio-temporal Transformer to fuse geometric and temporal cues across views, yielding a reliable initialization. Subsequently, we introduce a physics-aware 3D Gaussian optimization mechanism that incorporates tetrahedral deformation constraints, collision detection, and appearance disentanglement to achieve fine-grained reconstruction. Extensive experiments on both public and in-house datasets demonstrate that our method produces robust, artifact-free, high-fidelity results, offering an efficient solution for automated 4D asset generation.
This work addresses the challenge of achieving high-precision closed-loop control in tendon-driven anthropomorphic hands, which typically lack direct joint angle sensing. To circumvent the need for joint encoders, the authors propose a framework for joint angle estimation and control that combines a high-degree-of-freedom kinematic model—formulated using the Denavit–Hartenberg convention—with tendon displacement measurements and a simplified tension model to infer joint angles. A Jacobian-based PI controller augmented with feedforward compensation is then designed for accurate gesture tracking. Validated through nonlinear optimization and MuJoCo simulations on an Anatomically Correct Biomechatronic Hand, the approach successfully reproduces complex gestures with high fidelity, significantly enhancing control performance while preserving mechanical compactness.