Score
Designs and trains representation-learning pipelines and multi-head models that produce shared embeddings from MRI scans optimized simultaneously for multiple tasks (for example segmentation, classification, or reconstruction used as an auxiliary objective). Builds task-guided loss functions, multi-task architectures, and evaluation protocols to encode MRI structural features in a way that balances task-specific performance and improves robustness and generalization across cohorts.
Medical multimodal learning faces challenges including scarcity of paired data and reliance on proprietary or pretrained encoders. To address these, this paper proposes a single-encoder, parameter-sharing framework that unifies text and imaging modalities within a shared Transformer architecture. It introduces learnable modality embeddings for adaptive representation learning and designs a cross-modal parameter-sharing mechanism coupled with a joint contrastive alignment loss to alleviate low-resource generalization bottlenecks. Crucially, the approach eliminates modality-specific encoders. Evaluated across multiple medical multimodal benchmarks, it achieves significant improvements in few-shot settings (<1k samples): average retrieval accuracy increases by 4.2%, and classification F1 score improves by 3.8%. The core contribution is the first lightweight, parameter-shared multimodal representation learning paradigm explicitly designed for low-resource medical scenarios.
This study addresses the limitation of existing deep learning models in neuroimaging, which are typically confined to single tasks and struggle with cross-task knowledge transfer. To overcome this, the authors propose GenFAR, a modular deep learning framework trained jointly on 17 cognitive, clinical, and diagnostic tasks using brain MRI data from 49,246 individuals across 11 cohorts. The framework leverages an innovative task-sequencing mechanism and a novel Donor Score metric to identify five pivotal source tasks that substantially enhance sample efficiency and performance on downstream tasks. The resulting general-purpose brain representations demonstrate strong clinical relevance, significantly improving model accuracy on unseen tasks while markedly reducing the required training sample size.
Analyzing subtle, spatially sparse, and viewpoint-dependent lesions in brain MRI remains challenging. Method: We introduce the largest publicly available paired MRI–clinical report dataset to date (80K samples, a 10× scale-up), and propose a multi-view alignment representation learning framework featuring: (i) a novel implicit query-feature matching mechanism; (ii) quality- and diversity-driven multi-view embedding alignment; and (iii) integration of 3D slice-level feature disentanglement/aggregation with a document-retrieval-inspired cross-modal pretraining paradigm. Contribution/Results: Our approach achieves state-of-the-art performance across both vision-language understanding and pure-vision medical tasks. We open-source the BRAT foundation model, enabling zero-shot transfer and clinical report generation.
This work addresses the challenge of scarce and costly annotations in medical imaging by proposing COJEPA, a novel framework that extends I-JEPA to 3D brain MRI for the first time, integrating contrastive learning with a joint-embedding predictive architecture. The method enhances self-supervised representation learning on unlabeled T1-weighted images by simultaneously optimizing local structural predictability and global representational discriminability through foreground-aware block masking, hierarchical convolutional block embeddings, and world-space sinusoidal positional encoding. Experimental results demonstrate strong performance: the model achieves a rank@1 retrieval accuracy of 0.84 in zero-shot monozygotic twin identification, a mean absolute error of 2.55 years in age regression on OpenBHB, and competitive whole-tumor segmentation Dice scores on BraTS, matching state-of-the-art supervised methods.
Existing MR reconstruction methods prioritize image fidelity while neglecting their impact on downstream tasks (e.g., segmentation, classification), leading to cascaded performance degradation due to error propagation and domain shift. To address this, we propose a continual learning–based reconstruction optimization framework tailored for sequential multi-task deployment. For the first time, we introduce replay-based continual learning into MR reconstruction fine-tuning: a replay buffer jointly optimizes k-space domain reconstruction and downstream task gradients via a multi-task weighted loss, effectively mitigating catastrophic forgetting. Our approach employs a single reconstruction network that concurrently adapts to multiple downstream tasks—preserving high performance across all tasks while eliminating cascade mismatch. Extensive experiments demonstrate that our method significantly outperforms the conventional dual-network paradigm relying on independent optimization.
This study addresses the scarcity of annotated medical images and the heavy reliance of deep learning on manual annotations by proposing a multi-task self-supervised pre-training framework. Methodologically, this work innovatively jointly optimizes voxel-level brain age prediction, a domain-specific task, with image inpainting, a general-purpose task, to learn complementary neuroimaging representations and construct a highly generalizable foundation model. Experimental results demonstrate that the proposed framework significantly outperforms both single-task and from-scratch training baselines on segmentation tasks involving conditions such as multiple sclerosis. By effectively leveraging unlabeled data through complementary pretext tasks, this approach substantially mitigates the bottleneck imposed by data scarcity in medical image analysis.
研究对比了解剖特征与学习特征在脑MRI分析中的效果,提出结合解剖信息预训练的新方法,提升了生物年龄估计的准确性。
This work addresses the challenges of small sample sizes, low-quality labels, and high dimensionality in fMRI data that often lead to model overfitting. To this end, the authors propose BrainSimSiam, a lightweight self-supervised representation learning framework. It innovatively employs a positive-pair-only Siamese network architecture combined with fMRI-specific data augmentation, feature disentanglement, and consistency constraints to learn task-agnostic, robust functional brain representations—without requiring large-scale pretraining or negative samples. Experimental results demonstrate that the learned representations significantly outperform fully supervised baselines across multiple downstream classification and regression tasks and approach the performance of large-scale pretrained models, thereby substantially reducing reliance on computational resources.
This work addresses the domain shift in magnetic resonance imaging (MRI) caused by variations in magnetic field strength by proposing a 3D unpaired cross-field-strength image translation framework that does not require paired multi-field-strength data. The method leverages field-strength-conditioned content–style disentanglement pretraining to separate anatomical structure from field-strength-dependent contrast features, and integrates an AdaIN-modulated decoder with multi-field-strength discriminators to enable controllable image generation across arbitrary field strengths within a unified model. Evaluated on the MRIxFields dataset encompassing five field strengths and three modalities, the approach demonstrates high-fidelity preservation of 3D anatomical structures in diverse translation tasks—including Any-to-7T, 0.1T-to-High, and Any-to-Any—significantly enhancing both translation fidelity and flexibility.
This study addresses the challenge of negative transfer in multitask learning with multimodal clinical data, which hinders effective modeling of related yet heterogeneous clinical outcomes. To overcome this limitation, the authors propose a unified Transformer-based multitask framework incorporating an Orthogonal Task Decomposition (OrthTD) mechanism. This approach explicitly decouples shared and task-specific subspaces at the representation level and enforces geometric orthogonality constraints to suppress redundancy and isolate task-unique signals. Evaluated on data from 12,430 surgical patients for predicting four distinct clinical outcomes, the model achieves an average AUC of 87.5% and AUPRC of 37.2%, significantly outperforming existing methods—particularly excelling in the detection of rare events.