🤖 AI Summary
This study addresses the absence of end-to-end models and in-the-wild benchmarks for metric-scale 3D hand reconstruction from first-person stereo vision. To this end, we propose an end-to-end hand reconstruction method tailored for wearable stereo cameras. Our approach integrates stereo geometric constraints with temporal reasoning, introduces a pseudo-label training pipeline, and constructs an in-the-wild dataset validated against motion-capture ground truth. Experimental results demonstrate that the proposed model achieves state-of-the-art accuracy while preserving true metric scale even under monocular input. Furthermore, it exhibits strong robustness to missing viewpoints, extreme illumination, occlusion, and motion blur, along with superior out-of-distribution generalization capabilities.
📝 Abstract
Human dexterity is guided by two eyes watching two hands: binocular vision supplies the metric 3D structure that fine-grained manipulation consumes. Egocentric stereo is therefore the natural perceptual interface for robots, AR, and VR-yet metric 3D hand reconstruction from this very signal still has neither an end-to-end model nor an in-the-wild benchmark. We propose ESTHER, a model whose stereo geometry, temporal reasoning, and output representation are designed for wearable egocentric stereo. It is trained on pseudo-labels from a calibrated labeling pipeline and in turn assembles our benchmark ESTHER3D, an egocentric stereo hand dataset pairing a large in-the-wild training set of model-generated labels with a motion capture test set of true metric ground truth. Experiments show state-of-the-art accu?racy, superior external generalization, and robustness to the missing views, dropped frames, and lighting and motion blur extremes of real egocentric capture that break existing meth?ods. This robustness runs deeper than graceful degradation: stereo guidance teaches the model to bind apparent hand scale to metric depth, so it not only adapts to different stereo rigs and modalities with minimal fine-tuning, but more strikingly preserves true metric scale even after collapsing to a single monocular view.