🤖 AI Summary
This study addresses the limited generalization of six-degree-of-freedom control in ultrasound visual servoing, caused by the absence of out-of-plane motion cues in 2D images. To overcome this, we propose an end-to-end image-to-motion reasoning framework that employs a DINOv3 encoder integrated with low-rank adaptation and a bidirectional relational module. The method combines supervised relative pose learning, reconstruction-guided closed-loop adaptation, and bounded residual correction to achieve closed-loop probe pose alignment without requiring anatomical priors. Experiments on both public and self-collected datasets validate the effectiveness of the proposed approach. Furthermore, real-robot deployments demonstrate successful dynamic forearm tracking and target view alignment, significantly enhancing cross-target generalizability.
📝 Abstract
Ultrasound visual servoing is essential for autonomous robotic ultrasound, yet 6-DoF probe control from 2D B-mode images remains challenging due to limited and ambiguous out-of-plane motion cues. Existing methods typically rely on anatomical priors or handcrafted visual features, limiting their generalizability across imaging targets. Inspired by trackerless 3D ultrasound reconstruction, we propose Recon2Servo, a visual servoing framework that learns image-to-motion inference directly from B-mode images for 6-DoF probe control. A DINOv3 encoder with low-rank adaptation and a bidirectional relation module estimate the relative probe pose between current and target images to guide iterative closed-loop target-view alignment. The framework combines supervised relative-pose learning, reconstruction-guided closed-loop adaptation, and bounded residual pose correction to improve motion inference during servoing. Evaluations on a public dataset and an in-house dataset collected from 12 healthy volunteers using different ultrasound systems demonstrate its effectiveness in reconstructed-volume servoing. Additional real-robot demonstrations of target-view alignment and dynamic tracking on a human forearm are provided in the supplementary video: https://youtu.be/qwsOdI-GMYk.