TomoTransformer: Towards a Foundation Model for CT Reconstruction

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the distribution shift problem in conventional CT reconstruction models, which necessitate retraining when projection counts, angles, or resolutions vary. We propose a Transformer-based foundation model for universal CT reconstruction that treats locally filtered projections as tokens and introduces an innovative backprojection space design. By leveraging self-attention to predict missing views, the method achieves a geometrically well-posed, detector-size-invariant representation, enabling zero-shot generalization across arbitrary input and target angular configurations. Trained on large-scale, multi-source datasets, the proposed model significantly outperforms existing methods on multiple benchmark datasets and demonstrates strong practical generalizability in real-world nanoscale brain imaging.
📝 Abstract
Supervised deep learning has advanced sparse-view tomographic reconstruction. However, conventional models, which typically map filtered back-projection (FBP) images or sinograms to clean reconstructions, are brittle under distribution shifts. Because they require retraining whenever projection counts and angles, detector resolutions, or data distributions change, their deployment in real-world applications remains limited. To address this, we introduce TomoTransformer, a transformer-based architecture that treats each \textit{local} filtered projection as an individual token and predicts missing views via self-attention. Crucially, TomoTransformer operates in a \emph{back-projection space} that separates projections across spatial locations, making view interpolation geometrically well-posed and invariant to detector size. This design yields a single foundation model that can process any number of input projections, at arbitrary angular locations and detector dimensions, and query any number of target angles without retraining. Trained on a large-scale dataset spanning diverse medical CT anatomies and natural images, TomoTransformer generalizes effectively across anatomies, materials, and resolutions. Extensive evaluations on several benchmark sparse-view datasets show that TomoTransformer significantly outperforms concurrent multi-purpose models like ViewTrans and matches or exceeds strong protocol-specific baselines, while remaining fully agnostic to the number of input and target projections. Furthermore, the model demonstrates robust zero-shot generalization on real experimental nanoscale brain data collected from an X-ray synchrotron, showcasing its practical utility for real-world applications.
Problem

Research questions and friction points this paper is trying to address.

CT reconstruction
sparse-view tomography
distribution shift
foundation model
zero-shot generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Foundation Model
CT Reconstruction
TomoTransformer
Self-Attention
Zero-Shot Generalization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
AmirEhsan Khorashadizadeh
AmirEhsan Khorashadizadeh
PSI & EPFL
Machine LearningComputational ImagingTomographyPhase Retrieval
B
Benjamín Béjar
Swiss Data Science Center (SDSC) in Paul Scherrer Institute (PSI), Villigen, Switzerland