Cross-Modality Structural Guidance in 3D Latent Diffusion for Robust FLAIR Super-Resolution

📅 2026-06-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of anatomical detail loss in low-resolution or anisotropic FLAIR MRI scans, which arises from limited acquisition time and often leads existing super-resolution methods to generate structural hallucinations. To mitigate this, the authors propose MR-DiffuSR, a framework that leverages high-resolution T1-weighted images as structural priors within a 3D latent diffusion model. The method employs a cross-modal structural Swin attention mechanism to disentangle anatomical structure from modality-specific contrast, combined with a multi-scale degradation strategy and a DINOv3-based perceptual loss to enhance high-frequency detail fidelity and robustness to slice thickness variations. Evaluated on the ADNI-4 dataset, MR-DiffuSR achieves a PSNR of 32.46 dB, SSIM of 0.97, and LPIPS of 0.07; notably, it maintains a white matter hyperintensity segmentation Dice score of 0.63 even under 10× downsampling, substantially outperforming current state-of-the-art approaches.
📝 Abstract
High-resolution (HR) MRI acquisition is often hampered by scan time constraints, resulting in anisotropic or low-resolution scans (e.g., thick-slice FLAIR) that limit diagnostic accuracy. While deep learning-based super-resolution (SR) methods show promise, they often hallucinate anatomical details, which can compromise brain structural integrity. To mitigate this limitation, we introduce MR-DiffuSR, a Multi-Resolution Diffusion-based Super-Resolution framework that incorporates HR T1w structural image priors to guide the restoration of thick-slice FLAIR scans and operates in the 3D latent space. Our architecture introduces cross-modality structural swin-attention, which derives structural attention maps from the HR T1w and applies them to the low-resolution FLAIR latent features. This design disentangles anatomical structure from modality-specific contrast, effectively preventing hallucinations. Furthermore, we employ a mixed-scale degradation strategy, training the model on a continuum of downsampling factors to ensure robustness to varying slice thicknesses, while optimizing with a DINOv3-based perceptual loss to preserve high-frequency semantic details. Evaluated on the ADNI-4 dataset, MR-DiffuSR surpasses both CNN and 2D diffusion approaches, achieving an average PSNR of 32.46dB, SSIM of 0.97, and LPIPS of 0.07 across all downsampling factors. In downstream white matter hyperintensity segmentation, our model demonstrates exceptional robustness. While baseline performance collapses at 10x down-sampling (Dice: 0.51), MR-DiffuSR maintains a Dice score of 0.63, preserving utility even at 7mm equivalent slice thickness.
Problem

Research questions and friction points this paper is trying to address.

FLAIR super-resolution
MRI hallucination
structural integrity
cross-modality guidance
low-resolution MRI
Innovation

Methods, ideas, or system contributions that make the work stand out.

cross-modality guidance
3D latent diffusion
structural attention
mixed-scale degradation
perceptual loss
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
H
Haoyu Lan
University of Southern California
J
Jiazhen Zhang
Yale University
J
John Onofrey
Yale University
Bino Varghese
Bino Varghese
Keck School of Medicine
AIRadiomicsImagingImage processingPreclinical
N
Nasim Sheikh-Bahaei
University of Southern California
A
Arthur W. Toga
University of Southern California
Jeiran Choupan
Jeiran Choupan
USC Stevens Neuroimaging and Informatics Institute, Keck School of Medicine University of Southern
NeuroimagingMachine LearningBrain DecodingfMRI