🤖 AI Summary
This work addresses the scarcity of high-quality probe trajectories in robotic neck ultrasound guidance and the consequent difficulty in building accurate simulators. To overcome this, the authors propose a two-stage learning framework: first, an action-conditioned implicit diffusion world model is trained to predict future ultrasound images; then, this model is frozen and used as a reward source to fine-tune a goal-conditioned temporal Transformer that generates sequential probe motions. This approach represents the first integration of an action-conditioned diffusion world model with a goal-directed temporal policy, enabling high-precision ultrasound navigation without real-world interaction. Real closed-loop experiments on a self-collected dataset demonstrate 70.0% and 65.0% success rates for carotid and thyroid guidance tasks, respectively, validating the effectiveness of the learned ultrasound dynamics for goal-oriented navigation.
📝 Abstract
We present an action-conditioned world model framework for goal plane probe guidance in robotic ultrasound, with a focus on neck ultrasound scanning. Autonomous ultrasound tasks often require large numbers of probe-motion trajectories for training, but collecting high-quality demonstrations is labor-intensive and explicit simulators are difficult to build because ultrasound appearance depends on contact, tissue deformation, and view-dependent acoustic artifacts. We address this problem with a two-stage model-based learning pipeline. First, a latent conditional diffusion world model predicts future ultrasound observations from recent context frames, probe motions and temporal offset. Second, a goal-conditioned temporal transformer predicts ordered probe motions and is fine-tuned using rewards from the frozen world model. Experiments on the self-collected dataset show that the world model preserves action-dependent anatomical structure on target-directed scans. In real-world closed loop experiments, the framework achieves success rates of 70.0\% for carotid guidance and 65.0\% for thyroid guidance. These results demonstrate the potential of learned ultrasound dynamics for training goal-directed robotic probe navigation.