Semantic RGB--Depth Based Surgical Skill Assessment in Microscopic Stereo Videos

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of three-dimensional spatial relationships and the sparsity and unreliability of stereo depth in microsurgical assessment by proposing a semantic RGB-Depth framework. The method employs a regression-based deep fusion strategy that integrates sparse stereo matching with dense monocular depth estimation to generate reliable geometric representations. Furthermore, a hierarchical attention network is designed to jointly model the semantic segmentation stream and depth information, thereby precisely encoding the spatiotemporal interaction patterns between surgical instruments and anatomical structures. Evaluated on an ex vivo laryngeal surgery dataset, the proposed approach achieves a skill classification F1 score of 0.938, significantly outperforming pure RGB baselines. These results validate the critical role of dense geometric information in enhancing the accuracy of surgical assessment.
📝 Abstract
Objective assessment of microsurgical technical skill is essential for competency-based training and quality assurance, yet existing video-based approaches predominantly rely on RGB images and therefore overlook the 3D spatial relationships that characterize instrument-anatomy interactions. Although stereo operating microscopes provide complementary depth information, conventional stereo matching algorithms can produce sparse and unreliable depth estimates under high-magnification imaging conditions, limiting their use for automated skill assessment. This work presents a semantic RGB-Depth framework for surgical skill assessment from microscopic stereo videos. A regression-based depth fusion method combines sparse metric stereo depth with dense monocular depth estimates to generate a dense geometric representation of the surgical scene. This representation is integrated with semantically decomposed RGB streams corresponding to individual surgical instruments and surrounding anatomy. A hierarchical attention architecture jointly encodes these streams to capture discriminative patterns of instrument use and instrument-anatomy interaction across surgeons at different training levels. The framework was evaluated on 33 ex vivo transoral microlaryngeal procedures performed by six surgeons, comprising attending surgeons and surgical residents, using leave-one-surgeon-out cross-validation. The proposed semantic RGB-Depth model achieved an F1 score of 0.938 for skill-level classification, compared with 0.696 for semantic RGB and 0.929 for semantic depth. These results suggest that geometric information can improve automated surgical skill assessment from microscopic stereo videos. The learned spatial, temporal, and semantic attention patterns also support qualitative examination of the scene regions, video segments, and semantic streams emphasized by the model.
Problem

Research questions and friction points this paper is trying to address.

surgical skill assessment
stereo videos
depth estimation
microsurgery
RGB-Depth
Innovation

Methods, ideas, or system contributions that make the work stand out.

Semantic RGB-Depth Fusion
Regression-based Depth Fusion
Hierarchical Attention Architecture
Surgical Skill Assessment
Microscopic Stereo Videos
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jecia Z. Y. Mao
Laboratory of Computational Sensing and Robotics, Johns Hopkins University, MD, USA
S
Sue M. Cho
Laboratory of Computational Sensing and Robotics, Johns Hopkins University, MD, USA
F
Francis X. Creighton
Department of Otolaryngology–Head and Neck Surgery, Johns Hopkins University School of Medicine, MD, USA
D
Deepa Galaiya
Department of Otolaryngology–Head and Neck Surgery, Johns Hopkins University School of Medicine, MD, USA
Russell H. Taylor
Russell H. Taylor
John C. Malone Professor of Computer Science, Johns Hopkins University
RoboticsMedical RoboticsComputer-Integrated SurgeryComputer-Assisted Surgery
Manish Sahu
Manish Sahu
Johns Hopkins University
Computer VisionMachine LearningMedical RoboticsComputer Assisted InterventionsMedical Image Analysis