๐ค AI Summary
Severe artifacts and low visual quality in intraoperative cone-beam CT (CBCT) limit the accuracy of synthetic CT (sCT) generation. To address this, we propose an end-to-end differentiable multimodal fusion framework that jointly optimizes differentiable image registration and generative modeling for cross-modal alignment between preoperative CT and intraoperative CBCT, while simultaneously generating high-fidelity sCTโeliminating the need for sequential registration or manual parameter tuning. Our key contribution is a robust joint optimization mechanism explicitly designed for low-SNR CBCT and moderate anatomical misalignment (e.g., 5โ10 mm initial displacement). Evaluated on 90 clinical scenarios, our method significantly outperforms state-of-the-art baselines on 79 quantitative metrics, with particularly pronounced gains under challenging conditions of low CBCT signal-to-noise ratio and substantial initial CT-CBCT misalignment. This work establishes a new paradigm for high-fidelity sCT generation to support precise radiotherapy guidance.
๐ Abstract
Cone-Beam Computed Tomography (CBCT) is widely used for intraoperative imaging due to its rapid acquisition and low radiation dose. However, CBCT images typically suffer from artifacts and lower visual quality compared to conventional Computed Tomography (CT). A promising solution is synthetic CT (sCT) generation, where CBCT volumes are translated into the CT domain. In this work, we enhance sCT generation through multimodal learning by jointly leveraging intraoperative CBCT and preoperative CT data. To overcome the inherent misalignment between modalities, we introduce an end-to-end learnable registration module within the sCT pipeline. This model is evaluated on a controlled synthetic dataset, allowing precise manipulation of data quality and alignment parameters. Further, we validate its robustness and generalizability on two real-world clinical datasets. Experimental results demonstrate that integrating registration in multimodal sCT generation improves sCT quality, outperforming baseline multimodal methods in 79 out of 90 evaluation settings. Notably, the improvement is most significant in cases where CBCT quality is low and the preoperative CT is moderately misaligned.