🤖 AI Summary
This work addresses the scarcity of readily obtainable dense spatial attention labels in laparoscopic surgery, where surgical intent is highly specialized and difficult to annotate explicitly. The authors propose DiffeoAfford, a novel framework that introduces, for the first time, the concept of action-anchored tissue affordance. By leveraging tissue deformation tracking under diffeomorphic constraints and instrument trajectory analysis, the framework automatically generates affordance heatmaps as visual attention supervision signals without requiring frame-by-frame manual annotation. A real-time attention prediction model trained with these signals powers AffordView, an automated camera composition system capable of proactively anticipating critical surgical regions. Evaluated on real surgical procedures, AffordView demonstrates strong alignment with expert annotations and intraoperative gaze patterns, significantly reducing surgeons’ cognitive load, as validated across subjective, physiological, and behavioral metrics.
📝 Abstract
Computational attention models could help surgeons manage the visual demands of laparoscopy, but they require dense spatial labels that are difficult to obtain because surgical intent is highly specialized and tacit. Here, we introduce DiffeoAfford, an action-grounded tissue affordance framework that retrospectively derives visual attention supervision from completed surgical procedures. By combining diffeomorphism-constrained tissue tracking with instrument trajectory analysis, DiffeoAfford generates affordance hotspot labels without manual per-frame annotation. A real-time prediction model trained on these labels anticipates relevant surgical regions and enables AffordView, an assistive auto-framing system for laparoscopic visualization. The proposed framework aligns with expert annotations and intraoperative surgeon gaze, and reduces surgeon cognitive workload during real-world evaluations using subjective, physiological, and behavioral measures.