Robust Surgical Robotic Instrument Tracking via Sequential Multi-Cue Fusion and Sim-to-Real Self-Training

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of complex real-world scenarios and prohibitive pose annotation costs in surgical robotic instrument tracking, where existing detectors exhibit insufficient cross-domain adaptability. We propose a tracker-guided Sim-to-Real self-training framework that fuses multimodal features—including keypoints, boundaries, and masks—and integrates an uncertainty-aware extended Kalman filter with a Rauch–Tung–Striebel smoother to refine pose estimation. Furthermore, we introduce projection-based pseudo-labeling to generate supervision signals, enabling a closed-loop self-training pipeline without manual precise annotations. Experimental results demonstrate that the proposed method surpasses existing approaches in both accuracy and inference speed on real-world videos, substantially reducing reliance on costly annotated real data.
📝 Abstract
Efficient and robust tracking of surgical robotic instruments is important for robot-assisted minimally invasive surgery, yet remains challenging due to the complexity of surgical scenes and the unconventional geometry of surgical instruments. Keypoint-based approaches are efficient, but their performance depends on reliable feature detection. Improving these detectors with real-world supervision is difficult because accurate real-world annotations are costly to obtain at scale. To address this limitation, we introduce a tracker-guided self-training framework that adapts a model pretrained on synthetic images to unlabeled real-world videos. Given measured robot joint states, an uncertainty-aware EKF recursively corrects the instrument pose and the observable joint angles by comparing projected model features with detected keypoints, shaft boundaries, and mask-derived cues. An RTS smoother subsequently refines the resulting trajectory, which is projected into pseudo-labels for fine-tuning the feature detector without laborious pose annotations. Experiments on real-world videos demonstrate consistent improvements from self-training across all evaluated keypoint metrics, and the resulting model outperforms prior approaches in both accuracy and runtime. The code and data will be released upon publication.
Problem

Research questions and friction points this paper is trying to address.

Surgical instrument tracking
Keypoint detection
Sim-to-real adaptation
Annotation scarcity
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sim-to-Real Self-Training
Sequential Multi-Cue Fusion
Uncertainty-aware EKF
Pseudo-labeling
Surgical Instrument Tracking
🔎 Similar Papers
No similar papers found.