SurgGraph: Quantitative Laparoscopic Video Understanding via Geometry-Grounded Scene Graphs

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究通过SurgGraph方法,利用几何基础场景图从腹腔镜视频中生成定量场景图,以解决现有模型在理解手术场景时的不足。
📝 Abstract
Surgical videos are a primary resource for teaching trainees anatomy, tool usage, and procedural skills. Yet learning from them at scale requires systems that understand surgical scenes. Existing approaches fall short: vision-language models lack fine-grained domain reasoning, task-specific models do not generalize, and prior scene graphs omit clinically meaningful detail. We present SurgGraph, a training-free pipeline that generates quantitative scene graphs from surgical videos. Operating on segmentation masks and depth maps, SurgGraph encodes each clinically meaningful relation (attachment, occlusion, separation, tool actions) as a <subject, verb, object, value> tuple whose numeric value quantifies the relation's extent over time. Technical evaluations show more precise scene understanding than state-of-the-art surgical VLM baselines. We then build SurgGraphQA, a proof-of-concept learning application that retrieves meaningful and boundary-case exemplars and generates visual explanations and feedback. A study with 17 medical students and 2 resident surgeons shows significant learning gains, demonstrating its educational value.
Problem

Research questions and friction points this paper is trying to address.

surgical videos
scene graphs
clinical detail
Innovation

Methods, ideas, or system contributions that make the work stand out.

quantitative scene graphs
geometry-grounded
training-free pipeline
surgical video understanding
💼 Related Jobs
No related jobs found.