Enhancing Synthetic CT from CBCT via Multimodal Fusion and End-To-End Registration

๐Ÿ“… 2025-07-08
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
Severe artifacts and low visual quality in intraoperative cone-beam CT (CBCT) limit the accuracy of synthetic CT (sCT) generation. To address this, we propose an end-to-end differentiable multimodal fusion framework that jointly optimizes differentiable image registration and generative modeling for cross-modal alignment between preoperative CT and intraoperative CBCT, while simultaneously generating high-fidelity sCTโ€”eliminating the need for sequential registration or manual parameter tuning. Our key contribution is a robust joint optimization mechanism explicitly designed for low-SNR CBCT and moderate anatomical misalignment (e.g., 5โ€“10 mm initial displacement). Evaluated on 90 clinical scenarios, our method significantly outperforms state-of-the-art baselines on 79 quantitative metrics, with particularly pronounced gains under challenging conditions of low CBCT signal-to-noise ratio and substantial initial CT-CBCT misalignment. This work establishes a new paradigm for high-fidelity sCT generation to support precise radiotherapy guidance.

Technology Category

Computer Vision: Multi-modal VisionIntelligent Robots: Multimodal Perception & Sensor FusionConstraint Satisfaction and Optimization: Satisfiability Modulo Theories

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGUser Modeling, Personalization and Recommendation: Federated recommendation systems and personalizationSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
๐Ÿ“ Abstract
Cone-Beam Computed Tomography (CBCT) is widely used for intraoperative imaging due to its rapid acquisition and low radiation dose. However, CBCT images typically suffer from artifacts and lower visual quality compared to conventional Computed Tomography (CT). A promising solution is synthetic CT (sCT) generation, where CBCT volumes are translated into the CT domain. In this work, we enhance sCT generation through multimodal learning by jointly leveraging intraoperative CBCT and preoperative CT data. To overcome the inherent misalignment between modalities, we introduce an end-to-end learnable registration module within the sCT pipeline. This model is evaluated on a controlled synthetic dataset, allowing precise manipulation of data quality and alignment parameters. Further, we validate its robustness and generalizability on two real-world clinical datasets. Experimental results demonstrate that integrating registration in multimodal sCT generation improves sCT quality, outperforming baseline multimodal methods in 79 out of 90 evaluation settings. Notably, the improvement is most significant in cases where CBCT quality is low and the preoperative CT is moderately misaligned.
Problem

Research questions and friction points this paper is trying to address.

Improving synthetic CT quality from CBCT images
Addressing misalignment between CBCT and CT data
Enhancing multimodal fusion for better sCT generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal fusion of CBCT and CT data
End-to-end learnable registration module
Improved sCT quality via integrated registration
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
M
Maximilian Tschuchnig
Salzburg University of Applied Sciences, Puch bei Hallein, Austria
L
Lukas Lamminger
MedPhoton GmbH, Salzburg, Austria
P
Philipp Steininger
MedPhoton GmbH, Salzburg, Austria
Michael Gadermayr
Michael Gadermayr
Salzburg University of Applied Sciences
Machine LearningComputer VisionMedical Image Analysis