π€ AI Summary
This study addresses the limitations of conventional markerless hand-eye calibration, where binary masks cause detail loss, optimization instability, and insufficient accuracy, by proposing an RGB-based differentiable rendering framework. Methodologically, it replaces binary masks with RGB gradients to preserve internal object contour details and introduces a mask-guided image translation strategy that leverages deep networks to consistently fuse color and geometric multimodal features. Experimental evaluations on a UR5e robotic platform demonstrate grasping and insertion success rates of 88.9% and 57.4%, respectively, outperforming state-of-the-art methods by 46.3 and 48.1 percentage points. These results validate the proposed frameworkβs superior precision and robustness in practical manipulation tasks.
π Abstract
Accurate hand-eye calibration is crucial for precision manipulation. Traditional methods rely on markers, with their precision dependent on marker accuracy and observability. In contrast, markerless methods, such as learning-based approaches, use deep neural networks to directly extract keypoints or features from images, enabling the computation of hand-eye transformation with a single image and without the need for physical markers. Recently, differentiable rendering-based methods for hand-eye calibration have leveraged physical models to render binary masks and compare them with observations, enabling hand-eye calibration without fiducial markers in the calibration stage and providing interpretable optimization. While the state-of-the-art differentiable rendering methods achieve remarkable accuracy, the use of binary masks can result in the loss of internal profile details, reducing precision. Additionally, these methods can also suffer from unstable optimization and local minima.
In this study, we propose a novel RGB-based differentiable rendering framework that provides richer geometric and appearance cues by incorporating color and mask geometric features, thereby improving calibration accuracy and optimization stability. Additionally, we propose a mask-guided image-to-image translation method to ensure explicit preservation of color and geometric consistency throughout the translation. Our approach is validated through both simulation and real-world experiments, with results demonstrating strong accuracy and robustness and clear improvements over existing differentiable rendering methods. Our method achieves a grasping success rate of 88.9% and insertion success rate of 57.4% on the UR5e real-world experiment, outperforming the state-of-the-art differentiable rendering hand-eye calibration method EasyHeC by 46.3 and 48.1 percentage points, respectively.