🤖 AI Summary
This study addresses the challenge of limited data availability in advanced semiconductor manufacturing, where acquiring transmission electron microscopy (TEM) images is costly and sample scarcity hinders machine learning applications requiring diverse datasets. To overcome this bottleneck, the authors propose a high-fidelity synthetic image generation method based on denoising diffusion probabilistic models (DDPMs). Leveraging only 15 real TEM images, their approach employs a progressive patch-based training strategy to scale from local regions to full-image synthesis. The framework integrates custom TrivialAugment augmentation, cross-process-domain transfer, classifier guidance, and RePaint-style inpainting to enhance structural fidelity and physical realism. Generated images achieve MS-SSIM scores exceeding 0.98 and are validated by domain experts as effective for downstream tasks such as defect detection, segmentation, and metrology, demonstrating significant performance even under extremely data-constrained conditions.
📝 Abstract
Advanced semiconductor nodes drastically increased demand for Transmission Electron Microscopy (TEM), yet destructive sample preparation, slow imaging and high costs severely limit the availability of diverse datasets needed for downstream machine learning (ML). Synthetic data generation is becoming essential, but current generative models often miss TEM-specific noise, structural detail, and stochastic variability crucial for evaluation. We present a Denoising Diffusion Probabilistic Model (DDPM) framework for synthetic TEM image generation under extreme data scarcity. A progressive patch-based training strategy scales from low-resolution patches to full images, enabling from-scratch training with only 15 samples. We integrate a custom TrivialAugment adaptation, cross-process domain transfer, classifier guidance, and RePaint-style inpainting, culminating in full-image generation that preserves global structural and spatial relationships in compliance with FAB metrology requirements. Beyond synthesis, we repurpose DDPM feature representations for segmentation, partitioning encoder feature maps to obtain coherent region masks. Our synthetic images achieve up to MS-SSIM > 0.98 and qualitative expert assessment consistent with structural similarity results, facilitating downstream ML training for defect detection, segmentation, and metrology while preserving statistical and physical realism.