🤖 AI Summary
This study addresses the challenge of cross-modal knowledge transfer between fetal MRI and ultrasound arising from the lack of paired data. To this end, we propose WASP, a framework that introduces a novel weak alignment mechanism based on entropy-regularized optimal transport. By leveraging clinical metadata such as gestational age and diagnostic planes to construct spatiotemporal pairs, WASP effectively maps MRI representations into the ultrasound latent space. The method operates in a plug-and-play manner without fine-tuning backbone networks, enabling seamless integration with pretrained foundation models such as USFM and SAM-Med2D. Experimental results demonstrate that WASP significantly reduces gestational age estimation errors and improves standard plane classification accuracy in scenarios where MRI data is unavailable during inference, thereby validating the effectiveness of lightweight cross-modal knowledge transfer.
📝 Abstract
Magnetic Resonance Imaging (MRI) is widely regarded as the optimal sensor for fetal brain analysis due to its superior soft-tissue contrast and anatomical detail. However, its high cost and operational burden make it invasive and difficult to obtain at scale. Ultrasound (US), in contrast, is cheap, safe, and routinely acquired, and as a result it has produced substantially larger datasets and a growing ecosystem of pretrained models. This asymmetry raises a natural question: Can we teach a US-only model to understand fetal MRI from only a limited set of examples? The standard recipe, training a foundation model on subject-to-subject paired MRI-US scans, is not viable since no such paired fetal dataset is publicly available. In this paper we address this gap with Weakly Aligned Spatiotemporal Pairs (WASP), a framework that formulates cross-modal correspondence as an entropic Optimal Transport problem driven by clinical metadata, in particular Gestational Age (GA) and diagnostic planes, enabling the fitting of a lightweight alignment module that lifts MRI representations into the US latent space, without fine-tuning the backbone. Empirically, WASP yields its largest gains when MRI is unseen by the model during pretraining (on USFM, GA estimation error drops from 21.9 to 17.4 days and standard plane classification accuracy climbs from 61.9% to 69.0%), while providing smaller, backbone-dependent refinements for backbones pretrained on both modalities (e.g., BioMedParse GA estimation error from 6.5 to 6.0 days and SAM-Med2d plane accuracy from 83.3% to 88.1%). Code is available at https://github.com/miccunifi/WASP.