multi-task mri representation learning

Designs and trains representation-learning pipelines and multi-head models that produce shared embeddings from MRI scans optimized simultaneously for multiple tasks (for example segmentation, classification, or reconstruction used as an auxiliary objective). Builds task-guided loss functions, multi-task architectures, and evaluation protocols to encode MRI structural features in a way that balances task-specific performance and improves robustness and generalization across cohorts.

multi-taskmrirepresentationlearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.52
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

A Shared Encoder Approach to Multimodal Representation Learning

Mar 03, 2025
SR
Shuvendu Roy
🏛️ Vector Institute | Queen's University | York University

Medical multimodal learning faces challenges including scarcity of paired data and reliance on proprietary or pretrained encoders. To address these, this paper proposes a single-encoder, parameter-sharing framework that unifies text and imaging modalities within a shared Transformer architecture. It introduces learnable modality embeddings for adaptive representation learning and designs a cross-modal parameter-sharing mechanism coupled with a joint contrastive alignment loss to alleviate low-resource generalization bottlenecks. Crucially, the approach eliminates modality-specific encoders. Evaluated across multiple medical multimodal benchmarks, it achieves significant improvements in few-shot settings (<1k samples): average retrieval accuracy increases by 4.2%, and classification F1 score improves by 3.8%. The core contribution is the first lightweight, parameter-shared multimodal representation learning paradigm explicitly designed for low-resource medical scenarios.

Addresses scarcity of paired multimodal data in medical domain.Improves generalization with limited training data in medical applications.Proposes shared encoder framework for multimodal representation learning.

This study addresses the limitation of existing deep learning models in neuroimaging, which are typically confined to single tasks and struggle with cross-task knowledge transfer. To overcome this, the authors propose GenFAR, a modular deep learning framework trained jointly on 17 cognitive, clinical, and diagnostic tasks using brain MRI data from 49,246 individuals across 11 cohorts. The framework leverages an innovative task-sequencing mechanism and a novel Donor Score metric to identify five pivotal source tasks that substantially enhance sample efficiency and performance on downstream tasks. The resulting general-purpose brain representations demonstrate strong clinical relevance, significantly improving model accuracy on unseen tasks while markedly reducing the required training sample size.

brain representationdeep learningknowledge transfer

brat: Aligned Multi-View Embeddings for Brain MRI Analysis

Dec 21, 2025
MK
Maxime Kayser
🏛️ Memorial Sloan Kettering Cancer Center | University of Oxford

Analyzing subtle, spatially sparse, and viewpoint-dependent lesions in brain MRI remains challenging. Method: We introduce the largest publicly available paired MRI–clinical report dataset to date (80K samples, a 10× scale-up), and propose a multi-view alignment representation learning framework featuring: (i) a novel implicit query-feature matching mechanism; (ii) quality- and diversity-driven multi-view embedding alignment; and (iii) integration of 3D slice-level feature disentanglement/aggregation with a document-retrieval-inspired cross-modal pretraining paradigm. Contribution/Results: Our approach achieves state-of-the-art performance across both vision-language understanding and pure-vision medical tasks. We open-source the BRAT foundation model, enabling zero-shot transfer and clinical report generation.

Addresses challenges of varied abnormalities in 3D MRI scansAligns MRI embeddings with clinical report featuresDevelops a multi-view framework for brain MRI analysis

This work addresses the challenge of scarce and costly annotations in medical imaging by proposing COJEPA, a novel framework that extends I-JEPA to 3D brain MRI for the first time, integrating contrastive learning with a joint-embedding predictive architecture. The method enhances self-supervised representation learning on unlabeled T1-weighted images by simultaneously optimizing local structural predictability and global representational discriminability through foreground-aware block masking, hierarchical convolutional block embeddings, and world-space sinusoidal positional encoding. Experimental results demonstrate strong performance: the model achieves a rank@1 retrieval accuracy of 0.84 in zero-shot monozygotic twin identification, a mean absolute error of 2.55 years in age regression on OpenBHB, and competitive whole-tumor segmentation Dice scores on BraTS, matching state-of-the-art supervised methods.

contrastive learningjoint-embedding predictionrepresentation learning

Existing MR reconstruction methods prioritize image fidelity while neglecting their impact on downstream tasks (e.g., segmentation, classification), leading to cascaded performance degradation due to error propagation and domain shift. To address this, we propose a continual learning–based reconstruction optimization framework tailored for sequential multi-task deployment. For the first time, we introduce replay-based continual learning into MR reconstruction fine-tuning: a replay buffer jointly optimizes k-space domain reconstruction and downstream task gradients via a multi-task weighted loss, effectively mitigating catastrophic forgetting. Our approach employs a single reconstruction network that concurrently adapts to multiple downstream tasks—preserving high performance across all tasks while eliminating cascade mismatch. Extensive experiments demonstrate that our method significantly outperforms the conventional dual-network paradigm relying on independent optimization.

Addresses performance degradation from separate network trainingOptimizes MR reconstruction for multiple downstream tasksUses continual learning to prevent catastrophic forgetting

Latest Papers

What's happening recently
View more

This study addresses the scarcity of annotated medical images and the heavy reliance of deep learning on manual annotations by proposing a multi-task self-supervised pre-training framework. Methodologically, this work innovatively jointly optimizes voxel-level brain age prediction, a domain-specific task, with image inpainting, a general-purpose task, to learn complementary neuroimaging representations and construct a highly generalizable foundation model. Experimental results demonstrate that the proposed framework significantly outperforms both single-task and from-scratch training baselines on segmentation tasks involving conditions such as multiple sclerosis. By effectively leveraging unlabeled data through complementary pretext tasks, this approach substantially mitigates the bottleneck imposed by data scarcity in medical image analysis.

Brain MRIMedical Image SegmentationPretext Tasks

This work addresses the challenges of small sample sizes, low-quality labels, and high dimensionality in fMRI data that often lead to model overfitting. To this end, the authors propose BrainSimSiam, a lightweight self-supervised representation learning framework. It innovatively employs a positive-pair-only Siamese network architecture combined with fMRI-specific data augmentation, feature disentanglement, and consistency constraints to learn task-agnostic, robust functional brain representations—without requiring large-scale pretraining or negative samples. Experimental results demonstrate that the learned representations significantly outperform fully supervised baselines across multiple downstream classification and regression tasks and approach the performance of large-scale pretrained models, thereby substantially reducing reliance on computational resources.

fMRIhigh dimensionalitylabel quality

This work addresses the domain shift in magnetic resonance imaging (MRI) caused by variations in magnetic field strength by proposing a 3D unpaired cross-field-strength image translation framework that does not require paired multi-field-strength data. The method leverages field-strength-conditioned content–style disentanglement pretraining to separate anatomical structure from field-strength-dependent contrast features, and integrates an AdaIN-modulated decoder with multi-field-strength discriminators to enable controllable image generation across arbitrary field strengths within a unified model. Evaluated on the MRIxFields dataset encompassing five field strengths and three modalities, the approach demonstrates high-fidelity preservation of 3D anatomical structures in diverse translation tasks—including Any-to-7T, 0.1T-to-High, and Any-to-Any—significantly enhancing both translation fidelity and flexibility.

anatomical preservationcross-field MRI translationdomain shift

This study addresses the challenge of negative transfer in multitask learning with multimodal clinical data, which hinders effective modeling of related yet heterogeneous clinical outcomes. To overcome this limitation, the authors propose a unified Transformer-based multitask framework incorporating an Orthogonal Task Decomposition (OrthTD) mechanism. This approach explicitly decouples shared and task-specific subspaces at the representation level and enforces geometric orthogonality constraints to suppress redundancy and isolate task-unique signals. Evaluated on data from 12,430 surgical patients for predicting four distinct clinical outcomes, the model achieves an average AUC of 87.5% and AUPRC of 37.2%, significantly outperforming existing methods—particularly excelling in the detection of rare events.

multi-task learningmultimodal clinical datarepresentation disentanglement

Hot Scholars

SN

Shuteng Niu

Departmen of Artificial Intelligence & Informatics, Mayo Clinic
Transfer LearningGraph Representation LearningBiomedical Informatics
QL

Qizhen Lan

UTHealth Houston
Computer VisionKnowledge DistillationObject detectionMedical Imaging
QT

Qing Tian

University of Alabama at Birmingham
Computer VisionMachine LearningDeep LearningAutonomous Driving
LZ

Lijing Zhu

University of Houston-Clear Lake
Deep LearningHuman-object Interaction DetectionKnowledge Graph EmbeddingContinue Learning
XX

Xi Xiao

Oak Ridge National Laboratory | University of Alabama at Birmingham
LLM / MLLM EfficiencyImage / Video GenerationImage / Video Understanding