🤖 AI Summary
This study addresses the challenges of complex real-world scenarios and prohibitive pose annotation costs in surgical robotic instrument tracking, where existing detectors exhibit insufficient cross-domain adaptability. We propose a tracker-guided Sim-to-Real self-training framework that fuses multimodal features—including keypoints, boundaries, and masks—and integrates an uncertainty-aware extended Kalman filter with a Rauch–Tung–Striebel smoother to refine pose estimation. Furthermore, we introduce projection-based pseudo-labeling to generate supervision signals, enabling a closed-loop self-training pipeline without manual precise annotations. Experimental results demonstrate that the proposed method surpasses existing approaches in both accuracy and inference speed on real-world videos, substantially reducing reliance on costly annotated real data.
📝 Abstract
Efficient and robust tracking of surgical robotic instruments is important for robot-assisted minimally invasive surgery, yet remains challenging due to the complexity of surgical scenes and the unconventional geometry of surgical instruments. Keypoint-based approaches are efficient, but their performance depends on reliable feature detection. Improving these detectors with real-world supervision is difficult because accurate real-world annotations are costly to obtain at scale. To address this limitation, we introduce a tracker-guided self-training framework that adapts a model pretrained on synthetic images to unlabeled real-world videos. Given measured robot joint states, an uncertainty-aware EKF recursively corrects the instrument pose and the observable joint angles by comparing projected model features with detected keypoints, shaft boundaries, and mask-derived cues. An RTS smoother subsequently refines the resulting trajectory, which is projected into pseudo-labels for fine-tuning the feature detector without laborious pose annotations. Experiments on real-world videos demonstrate consistent improvements from self-training across all evaluated keypoint metrics, and the resulting model outperforms prior approaches in both accuracy and runtime. The code and data will be released upon publication.