Estimating Accurate Hand Pose in Camera Space with Vision Transformer

📅 2026-09-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种基于视觉变换器的方法,通过创新的监督机制和信息嵌入技术解决了单目RGB手部姿态估计中的深度模糊和局部-全局耦合问题。
📝 Abstract
Monocular RGB-based hand pose estimation has emerged as a critical research frontier in computer vision. The local hand pose estimation methods predict hand poses relative to the wrist, while global hand pose estimation also requires estimating the wrist's position in the camera coordinate system. However, this camera-space estimation confronts two fundamental challenges: (1) depth ambiguity in monocular settings, and (2) the coupling effect of hand local poses and global wrist positions in the perspective projections. In particular, this coupling reflects that the projections are jointly determined by local hand poses, wrist positions, and camera intrinsics. To overcome these challenges, our framework proposes two key innovations: Transformation-Isomorphism Supervision for hand-depth information extraction and Perspective Information Embedding for resolving above coupling effect of local pose and wrist position, both integrated within the mainstream encoder-decoder architecture. Besides, we propose a novel framerate-aware multi-dataset training strategy for sequential pose refinement. Our fully integrated approach achieves at most 37.1\% superiority in CS-MJE over SOTA on HO3D. Project page: https://github.com/Mine268/CS-ViT.
Problem

Research questions and friction points this paper is trying to address.

hand pose estimation
depth ambiguity
coupling effect
camera space
Innovation

Methods, ideas, or system contributions that make the work stand out.

Transformation-Isomorphism Supervision
Perspective Information Embedding
Vision Transformer
hand pose estimation
multi-dataset training
🔎 Similar Papers
K
Kaiwen Ren
Institute of Computing Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences, Beijing, China
Y
Yiran Jiang
Institute of Computing Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences, Beijing, China
Y
Yongjing Ye
Institute of Computing Technology, Chinese Academy of Sciences, Beijing, China
S
Shihong Xia
Institute of Computing Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences, Beijing, China