🤖 AI Summary
This study addresses the scarcity of surgical robot action data and the difficulty of determining tool contact timing from video streams by proposing an object-centric 3D contact flow framework that requires no action labels. The method integrates 3D tracking, flow matching generation, and proximity-based target extraction to predict trajectories and contact scores directly from stereo videos. This enables precise spatiotemporal modeling of tool-tissue interactions and optimized end-effector motion, facilitating embodiment-agnostic video utilization. Experimental results demonstrate that the proposed framework successfully completes 37 out of 39 tasks on a dVRK platform, significantly outperforming baseline methods. Furthermore, it achieves zero-shot transfer success rates of 85% and 70% on a humanoid laparoscopic robot, highlighting its strong generalization capabilities across distinct robotic embodiments.
📝 Abstract
Paired video-action demonstrations enable autonomous surgical behavior, but such data is scarce: robots perform roughly 1% of surgeries, while video-only data is abundant. Learning 3D object flow offers an embodiment-agnostic way to utilize video data, but flow alone specifies how an object should move, not where and when the tool should engage it, a distinction that is critical in surgery. We introduce SurgFlow, a framework that learns 3D Object-Centric Contact Flow from stereo surgical video without action labels. For each object point, it predicts a future 3D trajectory and contact scores. We extract targets via 3D tracking and tool-object proximity, train a flow matching generator to predict them, and use predicted contact to trigger grasp and release while optimizing end effector motion from flow. On the da Vinci Research Kit (dVRK), SurgFlow succeeds in 37 of 39 stage evaluations across tissue retraction, bimanual reveal, needle pickup, and handover, outperforming baselines trained on equal data with or without action labels. Zero-shot transfer to a humanoid-based laparoscopic robot achieves 85% and 70% average success under similar and novel camera viewpoints, respectively.