DRHeC: Differentiable Rendering for Hand-Eye Calibration with RGB-Based Gradients

πŸ“… 2026-09-29
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitations of conventional markerless hand-eye calibration, where binary masks cause detail loss, optimization instability, and insufficient accuracy, by proposing an RGB-based differentiable rendering framework. Methodologically, it replaces binary masks with RGB gradients to preserve internal object contour details and introduces a mask-guided image translation strategy that leverages deep networks to consistently fuse color and geometric multimodal features. Experimental evaluations on a UR5e robotic platform demonstrate grasping and insertion success rates of 88.9% and 57.4%, respectively, outperforming state-of-the-art methods by 46.3 and 48.1 percentage points. These results validate the proposed framework’s superior precision and robustness in practical manipulation tasks.
πŸ“ Abstract
Accurate hand-eye calibration is crucial for precision manipulation. Traditional methods rely on markers, with their precision dependent on marker accuracy and observability. In contrast, markerless methods, such as learning-based approaches, use deep neural networks to directly extract keypoints or features from images, enabling the computation of hand-eye transformation with a single image and without the need for physical markers. Recently, differentiable rendering-based methods for hand-eye calibration have leveraged physical models to render binary masks and compare them with observations, enabling hand-eye calibration without fiducial markers in the calibration stage and providing interpretable optimization. While the state-of-the-art differentiable rendering methods achieve remarkable accuracy, the use of binary masks can result in the loss of internal profile details, reducing precision. Additionally, these methods can also suffer from unstable optimization and local minima. In this study, we propose a novel RGB-based differentiable rendering framework that provides richer geometric and appearance cues by incorporating color and mask geometric features, thereby improving calibration accuracy and optimization stability. Additionally, we propose a mask-guided image-to-image translation method to ensure explicit preservation of color and geometric consistency throughout the translation. Our approach is validated through both simulation and real-world experiments, with results demonstrating strong accuracy and robustness and clear improvements over existing differentiable rendering methods. Our method achieves a grasping success rate of 88.9% and insertion success rate of 57.4% on the UR5e real-world experiment, outperforming the state-of-the-art differentiable rendering hand-eye calibration method EasyHeC by 46.3 and 48.1 percentage points, respectively.
Problem

Research questions and friction points this paper is trying to address.

Hand-Eye Calibration
Differentiable Rendering
Binary Mask Limitations
Optimization Stability
Local Minima
Innovation

Methods, ideas, or system contributions that make the work stand out.

Differentiable Rendering
Hand-Eye Calibration
RGB-Based Gradients
Image-to-Image Translation
Markerless
πŸ”Ž Similar Papers
No similar papers found.
X
Xiaotian Zhang
Department of Precision Engineering, School of Engineering, The University of Tokyo, Tokyo, Japan
Yusheng Wang
Yusheng Wang
The University of Tokyo
SLAMunderwater robotics3D image processing
N
Naoya Kagawa
FA Products Business Unit, DENSO WAVE INCORPORATED, Aichi, Japan
N
Noritaka Takamura
FA Products Business Unit, DENSO WAVE INCORPORATED, Aichi, Japan
K
Keiji Okuhara
FA Products Business Unit, DENSO WAVE INCORPORATED, Aichi, Japan
H
Hiroyasu Baba
FA Products Business Unit, DENSO WAVE INCORPORATED, Aichi, Japan
Jun Ota
Jun Ota
Research into Artifacts, Center for Engineering (RACE), School of Engg., The University of Tokyo
RoboticsProduction Engineering