Healthy Counterfactual Generation via Diffusion Inpainting for Mammography Classification

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the high false-negative rate in mammographic computer-aided diagnosis (CAD) systems, which frequently leads to clinical missed diagnoses. To mitigate this issue, we propose a diffusion model-based lesion erasure technique leveraging the RePaint inpainting strategy. This approach generates realistic healthy counterfactual images that preserve underlying anatomical structures, thereby enriching the training distribution. The augmented data is subsequently integrated with advanced architectures, including ConvNeXt and Vision Transformers, to optimize classifier performance. As a key contribution, this work pioneers the application of RePaint for lesion erasure and data augmentation in medical imaging. Extensive evaluations on the VinDr-Mammo dataset demonstrate that the proposed method significantly improves the sensitivity of four distinct classification models. Notably, substantial performance gains are observed under high-specificity thresholds, offering an effective solution for reducing missed diagnoses in clinical practice.
📝 Abstract
False negatives remain a critical limitation of computer-aided diagnosis (CAD) systems for breast cancer screening due to delayed detection and treatment. To address this issue, we propose a counterfactual data augmentation strategy that generates healthy mammograms by "erasing" lesions from anomalous images, thereby enriching the training distribution. We train a Denoising Diffusion Probabilistic Model on BI-RADS 1 (healthy) mammograms and use a RePaint-based sampling strategy to inpaint realistic normal tissue within annotated lesion bounding boxes. The resulting healthy counterfactuals replace annotated lesion regions with realistic healthy tissue while preserving patient-specific anatomical structure, as supported by similarity metrics between real and generated images. Image realism was further assessed by radiologists and found to be consistent with the original dataset quality. We evaluate counterfactual augmentation across four representative classifier architectures: a convolutional neural network (ConvNeXt), a vision transformer (ViT), a vision-language model pre-trained on mammogram-report pairs (Mammo-CLIP) and a multi-scale attention-based multiple-instance learning framework (FPN-MIL). Experiments conducted on the VinDr-Mammo dataset show improvements in sensitivity across all architectures, particularly at 80\% fixed specificity, contributing towards more reliable CAD systems for breast cancer. Code is available at: https://github.com/ines03garcia/diffusion-based-counterfactual-generation.
Problem

Research questions and friction points this paper is trying to address.

breast cancer screening
computer-aided diagnosis
false negatives
mammography classification
Innovation

Methods, ideas, or system contributions that make the work stand out.

Counterfactual Generation
Diffusion Inpainting
Mammography Classification
Data Augmentation
RePaint
🔎 Similar Papers
No similar papers found.
I
Inês Cruchinho Garcia
Institute for Systems and Robotics, Instituto Superior Técnico, Lisbon, Portugal
M
Mariana Mourão
Institute for Systems and Robotics, Instituto Superior Técnico, Lisbon, Portugal
F
Francisco Maria Calisto
Institute for Systems and Robotics, Instituto Superior Técnico, Lisbon, Portugal
Carlos Santiago
Carlos Santiago
Institute for Systems and Robotics (ISR/IST), LARSyS, Instituto Superior Técnico, Univ Lisboa
Computer VisionMachine learningDeep Learning
J
Jacinto Nascimento
Institute for Systems and Robotics, Instituto Superior Técnico, Lisbon, Portugal