Towards Better Cephalometric Landmark Detection with Diffusion Data Generation.

📅 2025-04-03
🏛️ IEEE Transactions on Medical Imaging
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Cephalometric landmark detection is hindered by high annotation costs and scarcity of high-quality labeled data, limiting deep learning model performance. To address this, we propose an anatomy-guided, text-conditioned diffusion generation framework—the first to fully automate both cephalometric X-ray image synthesis and precise landmark annotation. Our method introduces an anatomy-constrained controllable diffusion paradigm and constructs the first fine-grained medical-semantic text–image paired dataset for cephalometric radiographs. We integrate anatomical modeling, text-conditioned Stable Diffusion, and a DINOv2+DETR-based detection pipeline into an end-to-end optimized training framework. When trained solely on generated data, our landmark detector achieves an SDR (Success Detection Rate) of 82.2%, representing a 6.5-percentage-point improvement over prior baselines. All code and datasets are publicly released.

Technology Category

Computer Vision: Diffusion Models for VisionNatural Language Processing: Code Generation / Program Synthesis from Natural LanguageMachine Learning: Deep Generative Models & Autoencoders

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsEconomics, Online Markets and Human Computation: Humans versus LLMs for data annotation and labelingWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
Cephalometric landmark detection is essential for orthodontic diagnostics and treatment planning. Nevertheless, the scarcity of samples in data collection and the extensive effort required for manual annotation have significantly impeded the availability of diverse datasets. This limitation has restricted the effectiveness of deep learning-based detection methods, particularly those based on large-scale vision models. To address these challenges, we have developed an innovative data generation method capable of producing diverse cephalometric X-ray images along with corresponding annotations without human intervention. To achieve this, our approach initiates by constructing new cephalometric landmark annotations using anatomical priors. Then, we employ a diffusion-based generator to create realistic X-ray images that correspond closely with these annotations. To achieve precise control in producing samples with different attributes, we introduce a novel prompt cephalometric X-ray image dataset. This dataset includes real cephalometric X-ray images and detailed medical text prompts describing the images. By leveraging these detailed prompts, our method improves the generation process to control different styles and attributes. Facilitated by the large, diverse generated data, we introduce large-scale vision detection models into the cephalometric landmark detection task to improve accuracy. Experimental results demonstrate that training with the generated data substantially enhances the performance. Compared to methods without using the generated data, our approach improves the Success Detection Rate (SDR) by 6.5%, attaining a notable 82.2%. All code and data are available at: https://um-lab.github.io/cepha-generation/.
Problem

Research questions and friction points this paper is trying to address.

Addressing scarcity of diverse cephalometric X-ray datasets
Automating generation of annotated cephalometric X-ray images
Improving landmark detection accuracy with large-scale vision models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Generates diverse X-ray images using diffusion models
Constructs landmark annotations with anatomical priors
Uses text prompts to control image attributes
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Dongqian Guo
Dongqian Guo
University of Macau
Computer visionMedical image analysis
Wencheng Han
Wencheng Han
University of Macau
Medical Image AnalysisDepth EstimationAutomatic DriveMachine Learning
P
Pang Lyu
Department of Orthopaedic Surgery, Zhongshan Hospital, Fudan University, Shanghai, 200032 China
Y
Yuxi Zhou
Department of Periodontology, Justus-Liebig-University of Giessen, Germany, the Department of Periodontics, Stomatological Hospital of Guangzhou Medical University, China
Jianbing Shen
Jianbing Shen
Professor, University of Macau
Computer VisionMedical Image AnalysisVision and LanguageSelf-Driving CarsAI in Healthcare