Score
Designs and implements diffusion-based generative models (e.g., DDPMs) to synthesize realistic transmission electron microscopy (TEM) images, modeling instrument-specific noise and structural detail while preserving global spatial and structural relationships. Builds training, sampling, and evaluation pipelines that produce statistically diverse samples and support conditional controls such as classifier guidance and inpainting.
This study addresses the challenge of limited data availability in advanced semiconductor manufacturing, where acquiring transmission electron microscopy (TEM) images is costly and sample scarcity hinders machine learning applications requiring diverse datasets. To overcome this bottleneck, the authors propose a high-fidelity synthetic image generation method based on denoising diffusion probabilistic models (DDPMs). Leveraging only 15 real TEM images, their approach employs a progressive patch-based training strategy to scale from local regions to full-image synthesis. The framework integrates custom TrivialAugment augmentation, cross-process-domain transfer, classifier guidance, and RePaint-style inpainting to enhance structural fidelity and physical realism. Generated images achieve MS-SSIM scores exceeding 0.98 and are validated by domain experts as effective for downstream tasks such as defect detection, segmentation, and metrology, demonstrating significant performance even under extremely data-constrained conditions.
Conditional image synthesis with diffusion models suffers from a lack of systematic understanding due to architectural complexity, task heterogeneity, and diverse conditional mechanisms. Method: This paper introduces the first unified taxonomy for conditional diffusion modeling, categorizing approaches by *where* conditioning is injected—either into the denoising network architecture or the sampling process—and formalizes three paradigmatic stages: training, reuse, and specialization. It further classifies six mainstream sampling-time conditioning strategies. Contributions: Based on a structured analysis of over 100 works, the paper establishes a comprehensive knowledge framework and open-sources an authoritative resource repository (GitHub Awesome-Conditional-Diffusion-Models). It identifies persistent bottlenecks—including limited generalization, inefficient inference, and coarse-grained control—and proposes principled directions toward scalable, modular, and fine-grained conditional modeling.
The field of diffusion models for visual generation lacks systematic, pedagogically structured educational resources. Method: This project develops a teaching-oriented unified framework targeting undergraduate and graduate students, systematically integrating foundational probabilistic modeling—包括 forward diffusion, reverse denoising, stochastic differential equation (SDE) solvers, and score matching—with state-of-the-art conditional image and video generation. The framework emphasizes structured exposition of modeling principles, training paradigms, and sampling mechanisms to establish a clear, reproducible conceptual foundation. Contribution/Results: It significantly lowers the entry barrier for learners and fills a critical gap in introductory, comprehensive tutorials on diffusion models. The framework has become a widely adopted pedagogical benchmark and cross-disciplinary reference for both diffusion model instruction and applied research.
Diffusion model (DM)-generated images are increasingly indistinguishable from authentic camera-captured images, posing challenges for reliable detection. Method: This paper proposes a lightweight, local-statistics-based detection method that addresses image spatial non-stationarity by extracting local gradient distributions and texture features. It combines handcrafted features with conventional classifiers (e.g., SVM, Random Forest), eliminating the need for deep learning training. Contribution/Results: We systematically demonstrate—for the first time—that local statistical features significantly outperform global statistics in detecting DM-generated imagery, effectively resolving the non-stationary modeling challenge. Our method achieves substantially higher detection accuracy across diverse DM architectures compared to state-of-the-art global-statistics approaches. Moreover, it exhibits strong robustness against common post-processing operations—including JPEG compression and bilinear resizing—while maintaining high computational efficiency and practical deployability.
Diffusion probabilistic models (DPMs) suffer from slow sampling and numerical divergence under high classifier-free guidance (CFG) scales. To address this, we propose a stable and efficient high-order ODE-guided sampling framework. Our method introduces three key innovations: (1) the first high-order explicit solver specifically designed for data-prediction DPMs; (2) a threshold-based distribution correction mechanism to suppress gradient explosion at large CFG scales; and (3) a multi-step adaptive step-size reduction strategy to ensure numerical stability. Evaluated on both pixel-space and latent-space DPMs, our approach generates high-fidelity samples in only 15–20 function evaluations—over 5× faster than DDIM—while maintaining robust convergence across CFG=15–25. The implementation is publicly available.
This study addresses the challenge of detecting and classifying defects in irradiated metallic alloys from transmission electron microscopy (TEM) images, which is hindered by the scarcity of high-quality annotated data. To overcome this limitation, the authors propose a generative data augmentation approach that requires no manual labeling. Specifically, they introduce a mask-conditioned latent diffusion model (LDM) capable of controllably generating realistic TEM images along with corresponding multi-class defect masks. These synthetic data are then used to train a Mask R-CNN for joint defect detection and classification. Experimental results demonstrate that, under few-shot settings with only 10–100 real annotated images, the proposed method improves the harmonic mean of F1 scores by up to 0.02, confirming its effectiveness and practical utility.
This work addresses the underutilization of vast archives of transmission electron microscopy (TEM) data and their associated instrument parameter metadata. The authors propose a novel unpaired, physics-aware style transfer method that, for the first time, aligns high-angle annular dark-field scanning TEM (HAADF-STEM) images with their automatically recorded acquisition metadata through contrastive learning. This approach constructs a joint embedding space between images and metadata and integrates a generative network to enable metadata-conditioned image style transfer and denoising. Evaluated on a dataset of 7,330 real HAADF-STEM images, the method effectively synthesizes image appearances corresponding to diverse instrument settings and significantly enhances image quality, establishing a new paradigm for reusing unpublished electron microscopy data.
This work addresses the high computational cost of traditional microstructure simulations and the inefficiency of existing generative models, which require large amounts of paired data under continuous processing parameters. The authors propose a continuous conditional denoising diffusion model augmented with a neighborhood-aware loss training strategy, classifier-free guidance, and implicit sampling. This approach enables efficient generation of high-fidelity microstructure images from limited process–microstructure data pairs. The method substantially improves data utilization efficiency and generation quality, successfully reproducing key physical characteristics of low-carbon steel across varying manganese contents—including phase morphology, grain size distribution, phase fraction, and interfacial area distribution—demonstrating its capability to capture essential microstructural features with minimal training data.
This work addresses the challenge of accurately identifying atomic positions in short-exposure, high-speed, high-resolution transmission electron microscopy (HRTEM) images, which are severely degraded by strong noise. To this end, the authors propose a guided denoising network that jointly leverages spatial-domain bias and frequency-band statistical features. The method integrates a spatial-bias-guided weighted convolution module with a frequency-band-guided weighted filtering mechanism, underpinned by HRTEM-specific noise modeling, a realistically synthesized noise dataset, and a dedicated denoising architecture. This approach represents the first effort to enable efficient denoising driven by combined statistics from both spatial and frequency domains. Experimental results demonstrate that the proposed method consistently outperforms state-of-the-art techniques on both synthetic and real HRTEM data, significantly improving the accuracy of downstream tasks such as atomic localization.
This study addresses the prohibitive multi-particle training overhead and the limited sampling flexibility imposed by global fixed scoring rules when extending Diffusion Distribution Models (DDMs) to modern image generation. We propose an efficient single-stage class-conditional generation method based on a DiT latent space architecture. By deferring particle computations to deeper Transformer layers, the approach eliminates linear overhead, while a time-dependent dynamic scoring rule schedule is designed to overcome traditional performance bottlenecks. The method supports training from scratch without requiring teacher models or self-distillation. On ImageNet-256, it achieves an FID of 4.48 with 4 steps and 2.38 with 50 steps, notably without FID degradation as sampling steps increase. Furthermore, the proposed framework successfully generalizes to text-to-image generation tasks.