🤖 AI Summary
This work addresses the challenge in multi-view cardiac MRI diagnosis where existing models often conflate view-specific anatomical variations with pathological features, leading to shortcut learning and poor generalization—particularly under limited data regimes. To mitigate this, the authors propose the MoViD framework, which employs a dual-branch architecture built upon ViT-MAE. It explicitly disentangles view and disease representations through supervised contrastive learning combined with gradient reversal layers. Additionally, MoViD introduces, for the first time, unsupervised inter-frame motion cues to localize cardiac regions of interest and incorporates a focal re-weighted contrastive loss to suppress background distractions. Evaluated on a private venous thrombosis dataset as well as the M&Ms and M&Ms2 benchmarks, MoViD significantly outperforms standard Transformer baselines and achieves performance on par with large-scale pretrained models in both classification and segmentation tasks, effectively enabling causal disentanglement of view and disease factors in multi-view cardiac MRI.
📝 Abstract
Multi-view cardiac magnetic resonance (CMR) imaging provides complementary anatomical information and is widely used for noninvasive disease assessment. Recent transformer-based models have demonstrated strong representation learning capabilities for CMR analysis; however, they typically learn unified latent embeddings that entangle view-specific anatomical variations with disease-related features. Such entanglement biases classifiers toward structural attributes rather than view-invariant pathological patterns. This issue is exacerbated in low-data regimes, particularly for underrepresented cardiac conditions, where limited samples increase the susceptibility to shortcut learning and view-dependent decision boundaries. To address this, we propose a Motion-Guided View--Disease Disentanglement framework MoViD built upon a ViT-MAE backbone. The model explicitly factorizes latent representations into view-specific and disease-discriminative components using dual-branch supervised contrastive objectives and a gradient-reversal adversarial constraint that minimizes disease leakage into the view embedding. Additionally, an annotation-free temporal motion feature, derived from inter-frame difference maps, is introduced to localize the beating heart region and suppress background artifacts. A focal reweighting mechanism is incorporated into the contrastive loss to mitigate class imbalance. We evaluate the framework on a private clinical venous thrombosis dataset and two public benchmarks (M&Ms, M&Ms2). Across disease classification and cardiac segmentation tasks, our approach consistently outperforms standard transformer baselines and demonstrates competitive performance against large-scale pretrained foundation models, validating the efficacy of structural disentanglement in medical image analysis.