🤖 AI Summary
This study addresses the lack of dynamic anatomical prediction in robotic ultrasound navigation by proposing SonoGraph-WM, a scene graph-based world model. The method introduces a novel action- and goal-conditioning mechanism that leverages a unified Transformer to jointly predict future states and poses, enabling the inference of anatomical changes without synthesizing images. Furthermore, CT supervision and surface constraints are incorporated to generate training data, significantly reducing annotation dependency. Coupled with a receding-horizon planner, the framework achieves closed-loop navigation based on imagined trajectories. Experimental results demonstrate that spatial relationship prediction attains an F1 score exceeding 93%, while navigation success rates for the gallbladder and pancreas reach 77.5% and 75%, respectively, alongside a 73.7% target view arrival rate.
📝 Abstract
Ultrasound (US) acquisition depends on the operator's ability to interpret anatomy and anticipate how the view will change with probe motion. Many robotic US navigation methods select actions without explicitly predicting these anatomical changes. We propose SonoGraph-WM, an action- and goal-conditioned world model for anticipatory probe navigation. The model represents anatomy as scene graphs (SGs), capturing visible structures, their geometry, and spatial relationships without synthesizing US images. Given a history of SGs and probe poses, a unified Transformer jointly predicts future SGs and poses. A receding-horizon planner recursively imagines candidate trajectories, selects the shortest predicted path reaching a goal graph, and follows it over a short execution horizon before replanning from new observations. To reduce reliance on tracked and anatomically annotated US sequences, we generate aligned SG--pose training data from computed tomography (CT) label maps along surface-constrained probe trajectories. On four held-out CT cases, spatial relation F1 remains above 93% over 20 prediction steps, and closed-loop navigation achieves 77.50% and 75.00% success for the gallbladder and pancreas, respectively, using annotation-derived SGs. In robot--phantom navigation experiments with label-map-derived SGs, the planner reached the target view in 73.7% of trials. These findings support CT-supervised anatomical world modeling for probe planning and highlight the importance of frequent observation updates for reliable navigation. Project Page: https://noseefood.github.io/us-sonograph-wm/