🤖 AI Summary
This study addresses the bottlenecks of ultrasound world models, namely their reliance on costly synchronized video-pose data and difficulty in modeling cross-sectional sampling geometry. To overcome these limitations, this work proposes an interactive ultrasound world model. Methodologically, it introduces an acoustic sampling map to unify the representation of probe poses and imaging parameters, combined with 3D anatomical mask trajectories to guide the synthesis of action-video pairs. By adapting video foundation models and employing self-distillation from clinical data, the approach achieves real-label-free action prediction. Experimental results demonstrate that, in closed-loop planning, the proposed method reduces target distance and orientation errors by 29% and 38%, respectively, compared to visual servoing, thereby significantly improving predictive fidelity.
📝 Abstract
World models can enable autonomous ultrasound scanning by predicting the outcomes of probe motions from local observations. Learning this action--observation relationship typically relies on synchronized video--pose pairs, which are costly to collect at scale and largely unavailable in routine clinical recordings. Reliable action following further requires modeling ultrasound's cross-sectional sampling geometry. We present UltraWorld, a self-distillation recipe that transfers priors from clinical ultrasound videos into interactive world models without real action annotations. Starting from clinical videos, we adapt a video foundation model into an ultrasound generator conditioned on reference images and anatomical masks. Anatomical masks sampled along programmable trajectories through 3D anatomy provide spatial guidance for synthesizing action--video pairs. We then use these synthetic pairs to self-distill the generator into a world model that predicts future observations from local observations and actions, without requiring anatomical masks or other 3D assets at inference time. To further improve action following, we introduce the Acoustic Sampling Map (AsMap), which represents probe poses and imaging settings as pixel-wise 3D sampling positions, beam directions, and depths. Experiments demonstrate improved prediction fidelity and action following. Across nine simulated closed-loop local planning episodes, UltraWorld reduces the mean final distance to the goal and orientation error by 29\% and 38\%, respectively, compared with visual servoing. Project Page: https://ultraworld-project.github.io/.