🤖 AI Summary
This work addresses the pose estimation bias in low-cost motion capture devices caused by user-specific, fit-related, and cross-session variations by proposing the YOCO framework. YOCO pioneers the reformulation of per-user calibration as lightweight feed-forward adaptation: a HyperNetwork leverages minimal paired data to directly predict LoRA parameter updates, enabling rapid, fine-tuning-free calibration while keeping the MANO model frozen. This approach effectively preserves geometric priors and eliminates the need for individual operator optimization. Augmented with synthetic drift training, the proposed method is validated on both InterHand2.6M and real-world datasets, demonstrating significant improvements in hand state estimation quality and dexterous teleoperation performance.
📝 Abstract
Dexterous teleoperation requires reliable human-hand state estimations. However, common low-cost motion-capture gloves and markerless trackers often exhibit biases that vary across users, glove fit, and recording sessions, degrading retargeting and demonstration quality. We present YOCO, a fast few-shot, fine-tuning-free calibration framework that corrects biased hand-pose streams from a small set of paired raw and target poses. Instead of optimizing a separate model for every operator or session, YOCO conditions a calibration HyperNet on the paired examples and predicts LoRA-style updates for a frozen MANO hand-estimation module, turning per-user calibration into a lightweight feed-forward adaptation step while preserving the geometric prior of MANO and the efficiency of a compact estimator. We train YOCO with synthetic drift augmentations on InterHand2.6M and evaluate on augmented InterHand sequences, offline real glove data, and dexterous teleoperation tasks. Across these settings, YOCO improves calibration efficiency, hand-state estimation quality and teleoperation performance compared with uncalibrated input and standard calibration baselines.