🤖 AI Summary
This work addresses the challenges of inaccurate recognition and poor temporal coherence in surgical phase recognition, which stem from local blurriness, transient noise, and insufficient exploitation of semantic information. To overcome these limitations, the authors propose a novel approach that integrates structured surgical semantic knowledge. Specifically, they construct a hierarchical textual semantic representation encompassing intra-phase descriptions, inter-phase transitions, and fine-grained semantic units. Furthermore, they introduce a transition-aware segment construction (TAS-Con) module and a calibration mechanism (TAS-Calib), enabling efficient semantic aggregation and optimized phase representation without resorting to dense frame-level vision–language fusion. Experimental results demonstrate that the proposed method significantly improves both accuracy and robustness on the Cholec80 and LCRS-100 datasets.
📝 Abstract
Surgical video phase recognition is a fundamental task in computer-assisted intervention, supporting workflow understanding, intraoperative guidance, and surgical quality assessment. Although recent visual-temporal models have achieved promising progress, accurate and temporally coherent phase recognition remains challenging due to local visual ambiguity, transient prediction noise, and insufficient use of procedural semantics. To address these challenges, we propose HTT-Net, a Hierarchical Text-guided Transition modeling Network for surgical video phase recognition. The key idea is to introduce structured surgical semantic knowledge into phase-aware segment construction and semantic refinement. Specifically, we construct a hierarchical surgical semantic memory with intra-phase descriptions, inter-phase transition descriptions, and fine-grained semantic units. Based on this memory, the proposed Transition-Aware Segment Construction (TAS-Con) organizes frame-level evidence into coherent segment representations and handles boundary clips with inter-phase transition descriptions. Furthermore, we introduce Transition-Aware Segment Calibration (TAS-Calib), which calibrates phase-aware segment representations through hierarchical surgical semantics and improves discrimination under visual ambiguity without dense frame-level vision-language fusion. Experiments on Cholec80 and LCRS-100 demonstrate the effectiveness of HTT-Net for robust surgical video phase recognition.