Score
Design and implement latent diffusion generative models that are explicitly conditioned on anatomical structures or masks, operating in a compressed latent space for computational efficiency while enforcing anatomical consistency; these models are built to produce or edit images or volumes under anatomical constraints and to preserve sharp boundaries (e.g., pathology margins) during conditional synthesis or segmentation-guided generation.
Conditional image synthesis with diffusion models suffers from a lack of systematic understanding due to architectural complexity, task heterogeneity, and diverse conditional mechanisms. Method: This paper introduces the first unified taxonomy for conditional diffusion modeling, categorizing approaches by *where* conditioning is injected—either into the denoising network architecture or the sampling process—and formalizes three paradigmatic stages: training, reuse, and specialization. It further classifies six mainstream sampling-time conditioning strategies. Contributions: Based on a structured analysis of over 100 works, the paper establishes a comprehensive knowledge framework and open-sources an authoritative resource repository (GitHub Awesome-Conditional-Diffusion-Models). It identifies persistent bottlenecks—including limited generalization, inefficient inference, and coarse-grained control—and proposes principled directions toward scalable, modular, and fine-grained conditional modeling.
This work addresses the scarcity of high-quality brain segmentation masks in non-contrast CT neuroimaging, particularly due to the high annotation cost and substantial variability of ischemic infarcts. To tackle this challenge, the authors propose an anatomy-preserving generative framework that, for the first time, integrates a diffusion model into the latent space of a variational autoencoder (VAE) trained on segmentation masks. The method enables unconditional generation of multi-class brain tissue masks containing ischemic infarcts and supports coarse control over lesion presence via binary prompts. By leveraging a frozen VAE decoder to reconstruct masks, the approach effectively preserves global anatomical structure, discrete semantic labels, and realistic pathological variations while avoiding common structural artifacts associated with pixel-level generative models.
Existing 3D anatomical generative models struggle to simultaneously achieve geometric controllability and anatomical fidelity. To address this, we propose an interpretable, programmable guidance framework based on ellipsoidal primitives. Our method decouples the modeling of size, shape, and position parameters for individual anatomical structures, integrates multi-tissue segmentation maps and geometric moment loss into an unconditional diffusion model, and injects 3D ellipsoidal guidance signals during the reverse diffusion process. This enables inference-time, independent or compositional geometric editing of multiple tissues—without retraining. Evaluated on complex multi-organ structures such as the heart, our approach achieves millimeter-level morphological editing precision while preserving high anatomical fidelity. It supports flexible, structured generation of anatomically plausible configurations, significantly enhancing controllability and interpretability in medical image synthesis.
To address class imbalance, privacy constraints, and insufficient coverage of modality–pathology combinations in clinical brain MRI data, this paper introduces the first latent diffusion generative model jointly controllable across multiple pathologies (healthy, glioblastoma, multiple sclerosis, dementia) and multiple MRI modalities (T1w, T1ce, T2w, FLAIR, PD). The method employs conditional embedding encoding to achieve fine-grained, disentangled control over pathology and modality—enabling zero-shot cross-configuration extrapolation to unseen modality–pathology pairings. Built upon the Latent Diffusion framework, the model synthesizes high-fidelity images, achieving significantly lower FID and higher MS-SSIM scores than baseline methods. It effectively augments rare-class samples, thereby enhancing the robustness and diagnostic reliability of downstream models. Crucially, the approach balances high-quality data augmentation with stringent patient privacy preservation, as no raw sensitive data is shared or stored during generation.
Conventional registration methods struggle with large anatomical variations across subpopulations (e.g., age- or disease-specific cohorts), while generative models often produce anatomically implausible artifacts. Method: We propose a latent diffusion model (LDM)-based framework that directly synthesizes nonlinear deformation fields—not images—thereby eliminating anatomical hallucination. Our approach integrates differentiable multi-scale registration with neighborhood consistency constraints to ensure structural plausibility and anatomical fidelity. Contribution/Results: Evaluated on 5,000 brain and whole-body MRI scans from UK Biobank, our method generates smooth, artifact-free, and highly generalizable population atlases. It significantly outperforms traditional registration and GAN-based approaches in both qualitative and quantitative assessments. The resulting deformation fields are interpretable, robust, and enable fine-grained analysis of anatomical differences—such as those associated with aging or pathological morphology—establishing a new paradigm for population-specific atlas construction.
To address the challenges of small lesion size, complex anatomical structures, and scarce annotated data in abdominal lymph node segmentation, this paper proposes LN-DDPM—a conditional diffusion model incorporating dual conditioning on global anatomical structure and local details—within a generation-segmentation joint pipeline: high-fidelity paired CT images and masks are first synthesized, then fed into nnU-Net for segmentation. Key innovations include multi-scale mask guidance and explicit anatomical prior embedding, enabling precise decoupled modeling of lymph nodes and surrounding tissues. On an abdominal lymph node dataset, LN-DDPM achieves significantly superior synthetic image quality over GAN- and VAE-based baselines. Downstream segmentation yields a 4.2% Dice score improvement and an 18.7% increase in small-object detection rate, demonstrating synergistic gains between generative fidelity and segmentation performance.
Medical images are frequently compromised by artifacts, missing regions, or pathological alterations, which can undermine diagnostic reliability. This work presents a systematic review of diffusion model–based approaches for medical image inpainting and introduces the first taxonomy specifically tailored to this domain. The proposed framework encompasses prevailing architectures—such as Denoising Diffusion Probabilistic Models (DDPM) and Latent Diffusion Models (LDM)—alongside key clinical applications (e.g., MRI and CT), benchmark datasets, and evaluation protocols. Empirical analysis demonstrates that diffusion models excel at generating anatomically plausible reconstructions, yet critical challenges persist, notably the absence of standardized benchmarks and limited data diversity. By synthesizing current advances and identifying open problems, this study offers a structured foundation to guide future research in medical image restoration.
This study addresses the challenge of classifying diffuse gliomas across multi-center MRI datasets, where domain shift and scarce annotations hinder performance. The authors propose an anatomy-guided latent diffusion model within a two-stage framework: first, a 3D variational autoencoder learns anatomical priors from the source domain; then, a tumor mask-conditioned latent diffusion model, steered by ControlNet, synthesizes structurally coherent and boundary-sharp 3D glioma MRIs using only 16 target-domain samples. This work is the first to integrate anatomical priors with ControlNet-guided diffusion for data-efficient, few-shot cross-domain medical image synthesis. Experiments demonstrate that the method achieves a Fréchet Inception Distance (FID) of 85.40 and a downstream classification AUC of 0.987 under extreme data scarcity, significantly outperforming GAN-based baselines.
研究使用图像条件扩散模型来检测头颈部CT中器官风险分割的错误,以提高放射治疗规划的质量保证。
This work addresses the limitation of existing diffusion models in high-dimensional generation, which often ignore the intrinsic manifold geometry of data, while conventional latent diffusion models impose an Euclidean structure that struggles to capture complex geometries under data sparsity. To overcome this, we propose the Intrinsic Latent Diffusion Model (ILDM), the first framework to integrate Riemannian manifold geometry into the diffusion process. ILDM treats the latent space as a coordinate chart of an unknown manifold and jointly models geometric structure and uncertainty via a probabilistic decoder. We introduce a Riemannian–Euclidean hybrid forward diffusion mechanism, supported by a local uncertainty–driven diffusion strategy, a probabilistic metric tensor, and a tailored approximate denoising score matching objective. Experiments on COIL-100, MNIST, and cardiac MRI demonstrate that ILDM significantly outperforms existing methods, achieving lower FID and LPIPS scores and superior generation quality.
This work addresses the challenge of achieving precise and controllable image generation with diffusion models in the absence of large-scale annotated data. The authors propose a self-conditioning mechanism leveraging pretrained self-supervised representations, which identifies semantic directions in the representation space to guide the diffusion process without requiring labeled conditions. This approach not only enhances unconditional generation quality but also constructs a smooth and disentangled controllable generation space. Experimental results demonstrate that the proposed method achieves superior performance in image generation and editing tasks, excelling in controllability, smoothness, and disentanglement compared to existing alternatives.