anatomy-conditioned latent diffusion

Design and implement latent diffusion generative models that are explicitly conditioned on anatomical structures or masks, operating in a compressed latent space for computational efficiency while enforcing anatomical consistency; these models are built to produce or edit images or volumes under anatomical constraints and to preserve sharp boundaries (e.g., pathology margins) during conditional synthesis or segmentation-guided generation.

anatomy-conditionedlatentdiffusion

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.59
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the scarcity of high-quality brain segmentation masks in non-contrast CT neuroimaging, particularly due to the high annotation cost and substantial variability of ischemic infarcts. To tackle this challenge, the authors propose an anatomy-preserving generative framework that, for the first time, integrates a diffusion model into the latent space of a variational autoencoder (VAE) trained on segmentation masks. The method enables unconditional generation of multi-class brain tissue masks containing ischemic infarcts and supports coarse control over lesion presence via binary prompts. By leveraging a frozen VAE decoder to reconstruct masks, the approach effectively preserves global anatomical structure, discrete semantic labels, and realistic pathological variations while avoiding common structural artifacts associated with pixel-level generative models.

data scarcityischemic infarctmedical image analysis

CardioComposer: Flexible and Compositional Anatomical Structure Generation with Disentangled Geometric Guidance

Sep 08, 2025
KK
Karim Kadry
🏛️ Massachusetts Institute of Technology | Brigham and Women’s Hospital | American University in Cairo

Existing 3D anatomical generative models struggle to simultaneously achieve geometric controllability and anatomical fidelity. To address this, we propose an interpretable, programmable guidance framework based on ellipsoidal primitives. Our method decouples the modeling of size, shape, and position parameters for individual anatomical structures, integrates multi-tissue segmentation maps and geometric moment loss into an unconditional diffusion model, and injects 3D ellipsoidal guidance signals during the reverse diffusion process. This enables inference-time, independent or compositional geometric editing of multiple tissues—without retraining. Evaluated on complex multi-organ structures such as the heart, our approach achieves millimeter-level morphological editing precision while preserving high anatomical fidelity. It supports flexible, structured generation of anatomically plausible configurations, significantly enhancing controllability and interpretability in medical image synthesis.

Balancing controllability and anatomical realism in 3D anatomy generationEnabling compositional multi-component constraints during anatomical inferenceProviding independent control over size, shape, and position of tissues

Diffusion Models for conditional MRI generation

Feb 25, 2025
MH
Miguel Herencia Garc'ia del Castillo
🏛️ Ainovis.health | Telefónica Innovación Digital

To address class imbalance, privacy constraints, and insufficient coverage of modality–pathology combinations in clinical brain MRI data, this paper introduces the first latent diffusion generative model jointly controllable across multiple pathologies (healthy, glioblastoma, multiple sclerosis, dementia) and multiple MRI modalities (T1w, T1ce, T2w, FLAIR, PD). The method employs conditional embedding encoding to achieve fine-grained, disentangled control over pathology and modality—enabling zero-shot cross-configuration extrapolation to unseen modality–pathology pairings. Built upon the Latent Diffusion framework, the model synthesizes high-fidelity images, achieving significantly lower FID and higher MS-SSIM scores than baseline methods. It effectively augments rare-class samples, thereby enhancing the robustness and diagnostic reliability of downstream models. Crucially, the approach balances high-quality data augmentation with stringent patient privacy preservation, as no raw sensitive data is shared or stored during generation.

Condition on pathology and modalityEnhance clinical dataset diversityGenerate brain MRI images

Diff-Def: Diffusion-Generated Deformation Fields for Conditional Atlases

Mar 25, 2024
SS
Sophie Starck
🏛️ Technical University of Munich | Imperial College London | FAU Erlangen-Nürnberg

Conventional registration methods struggle with large anatomical variations across subpopulations (e.g., age- or disease-specific cohorts), while generative models often produce anatomically implausible artifacts. Method: We propose a latent diffusion model (LDM)-based framework that directly synthesizes nonlinear deformation fields—not images—thereby eliminating anatomical hallucination. Our approach integrates differentiable multi-scale registration with neighborhood consistency constraints to ensure structural plausibility and anatomical fidelity. Contribution/Results: Evaluated on 5,000 brain and whole-body MRI scans from UK Biobank, our method generates smooth, artifact-free, and highly generalizable population atlases. It significantly outperforms traditional registration and GAN-based approaches in both qualitative and quantitative assessments. The resulting deformation fields are interpretable, robust, and enable fine-grained analysis of anatomical differences—such as those associated with aging or pathological morphology—establishing a new paradigm for population-specific atlas construction.

Avoids hallucinations and ensures structural integrity in atlas generationGenerates deformation fields for conditional atlases using diffusion modelsHandles large anatomical variations better than registration-based methods

CT Synthesis with Conditional Diffusion Models for Abdominal Lymph Node Segmentation

Mar 26, 2024
YY
Yongrui Yu
🏛️ Shanghai Jiao Tong University | The First Hospital of China Medical University | Shanghai AI Laboratory

To address the challenges of small lesion size, complex anatomical structures, and scarce annotated data in abdominal lymph node segmentation, this paper proposes LN-DDPM—a conditional diffusion model incorporating dual conditioning on global anatomical structure and local details—within a generation-segmentation joint pipeline: high-fidelity paired CT images and masks are first synthesized, then fed into nnU-Net for segmentation. Key innovations include multi-scale mask guidance and explicit anatomical prior embedding, enabling precise decoupled modeling of lymph nodes and surrounding tissues. On an abdominal lymph node dataset, LN-DDPM achieves significantly superior synthetic image quality over GAN- and VAE-based baselines. Downstream segmentation yields a 4.2% Dice score improvement and an 18.7% increase in small-object detection rate, demonstrating synergistic gains between generative fidelity and segmentation performance.

Addresses limited annotated data and complex abdominal environment challengesEnhances segmentation accuracy using conditional diffusion models and nnU-NetImproves abdominal lymph node segmentation via synthetic data generation

Latest Papers

What's happening recently
View more

Medical images are frequently compromised by artifacts, missing regions, or pathological alterations, which can undermine diagnostic reliability. This work presents a systematic review of diffusion model–based approaches for medical image inpainting and introduces the first taxonomy specifically tailored to this domain. The proposed framework encompasses prevailing architectures—such as Denoising Diffusion Probabilistic Models (DDPM) and Latent Diffusion Models (LDM)—alongside key clinical applications (e.g., MRI and CT), benchmark datasets, and evaluation protocols. Empirical analysis demonstrates that diffusion models excel at generating anatomically plausible reconstructions, yet critical challenges persist, notably the absence of standardized benchmarks and limited data diversity. By synthesizing current advances and identifying open problems, this study offers a structured foundation to guide future research in medical image restoration.

anatomical consistencyartifact removaldiagnostic reliability

This study addresses the challenge of classifying diffuse gliomas across multi-center MRI datasets, where domain shift and scarce annotations hinder performance. The authors propose an anatomy-guided latent diffusion model within a two-stage framework: first, a 3D variational autoencoder learns anatomical priors from the source domain; then, a tumor mask-conditioned latent diffusion model, steered by ControlNet, synthesizes structurally coherent and boundary-sharp 3D glioma MRIs using only 16 target-domain samples. This work is the first to integrate anatomical priors with ControlNet-guided diffusion for data-efficient, few-shot cross-domain medical image synthesis. Experiments demonstrate that the method achieves a Fréchet Inception Distance (FID) of 85.40 and a downstream classification AUC of 0.987 under extreme data scarcity, significantly outperforming GAN-based baselines.

3D MRI synthesisdata scarcitydomain shift

This work addresses the limitation of existing diffusion models in high-dimensional generation, which often ignore the intrinsic manifold geometry of data, while conventional latent diffusion models impose an Euclidean structure that struggles to capture complex geometries under data sparsity. To overcome this, we propose the Intrinsic Latent Diffusion Model (ILDM), the first framework to integrate Riemannian manifold geometry into the diffusion process. ILDM treats the latent space as a coordinate chart of an unknown manifold and jointly models geometric structure and uncertainty via a probabilistic decoder. We introduce a Riemannian–Euclidean hybrid forward diffusion mechanism, supported by a local uncertainty–driven diffusion strategy, a probabilistic metric tensor, and a tailored approximate denoising score matching objective. Experiments on COIL-100, MNIST, and cardiac MRI demonstrate that ILDM significantly outperforms existing methods, achieving lower FID and LPIPS scores and superior generation quality.

diffusion modelslatent spacemanifold geometry

This work addresses the challenge of achieving precise and controllable image generation with diffusion models in the absence of large-scale annotated data. The authors propose a self-conditioning mechanism leveraging pretrained self-supervised representations, which identifies semantic directions in the representation space to guide the diffusion process without requiring labeled conditions. This approach not only enhances unconditional generation quality but also constructs a smooth and disentangled controllable generation space. Experimental results demonstrate that the proposed method achieves superior performance in image generation and editing tasks, excelling in controllability, smoothness, and disentanglement compared to existing alternatives.

controllable image generationdiffusion modelsgeneration control

Hot Scholars

LJ

Liming Jiang

Senior Research Scientist, ByteDance / TikTok, USA
Computer VisionGenerative AI
GL

Guixu Lin

The University of Tokyo
Video GenerationEvent Camera based VisionComputer Vision
SH

Shengfeng He

Singapore Management University
Visual ComputingGenerative ModelsComputer VisionComputational Photography
WS

Wei Song

Zhejiang University, Westlake University, Shanghai Innovation Institute
Artificial IntelligenceMulti-modal LearningMLLMs
LL

Linjie Luo

Research Manager at ByteDance AI Lab
Computer GraphicsComputer Vision