UltraWorld: Learning Interactive Ultrasound World Models from Untracked Clinical Videos with Acoustic Sampling Map

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the bottlenecks of ultrasound world models, namely their reliance on costly synchronized video-pose data and difficulty in modeling cross-sectional sampling geometry. To overcome these limitations, this work proposes an interactive ultrasound world model. Methodologically, it introduces an acoustic sampling map to unify the representation of probe poses and imaging parameters, combined with 3D anatomical mask trajectories to guide the synthesis of action-video pairs. By adapting video foundation models and employing self-distillation from clinical data, the approach achieves real-label-free action prediction. Experimental results demonstrate that, in closed-loop planning, the proposed method reduces target distance and orientation errors by 29% and 38%, respectively, compared to visual servoing, thereby significantly improving predictive fidelity.
📝 Abstract
World models can enable autonomous ultrasound scanning by predicting the outcomes of probe motions from local observations. Learning this action--observation relationship typically relies on synchronized video--pose pairs, which are costly to collect at scale and largely unavailable in routine clinical recordings. Reliable action following further requires modeling ultrasound's cross-sectional sampling geometry. We present UltraWorld, a self-distillation recipe that transfers priors from clinical ultrasound videos into interactive world models without real action annotations. Starting from clinical videos, we adapt a video foundation model into an ultrasound generator conditioned on reference images and anatomical masks. Anatomical masks sampled along programmable trajectories through 3D anatomy provide spatial guidance for synthesizing action--video pairs. We then use these synthetic pairs to self-distill the generator into a world model that predicts future observations from local observations and actions, without requiring anatomical masks or other 3D assets at inference time. To further improve action following, we introduce the Acoustic Sampling Map (AsMap), which represents probe poses and imaging settings as pixel-wise 3D sampling positions, beam directions, and depths. Experiments demonstrate improved prediction fidelity and action following. Across nine simulated closed-loop local planning episodes, UltraWorld reduces the mean final distance to the goal and orientation error by 29\% and 38\%, respectively, compared with visual servoing. Project Page: https://ultraworld-project.github.io/.
Problem

Research questions and friction points this paper is trying to address.

World Models
Autonomous Ultrasound Scanning
Untracked Clinical Videos
Action Following
Ultrasound Sampling Geometry
Innovation

Methods, ideas, or system contributions that make the work stand out.

World Models
Self-Distillation
Ultrasound
Acoustic Sampling Map
Video Foundation Model
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
K
Keke Yang
The Chinese University of Hong Kong
E
Erqi Wang
The Chinese University of Hong Kong
S
Sainan Guan
The Eighth Affiliated Hospital, Sun Yat-sen University
Hongliang Ren
Hongliang Ren
Chinese University of Hong Kong | National University of Singapore | JHU/Harvard(RF) | CUHK(PhD)
Biorobotics & intelligent systemsmedical mechatronicscontinuumsoft flexible robots/sensorsmultisensory perception