Representation-driven Endoscopic Visual Embedding Alignment for Latent Generation

📅 2026-08-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenges of large distributional discrepancies between natural and clinical endoscopic images and the high computational cost of training large diffusion models by proposing REVEAL, a novel approach that introduces representation alignment into endoscopic image generation for the first time. Leveraging a GastroNet-5M–pretrained domain-specific encoder, REVEAL aligns the diffusion model’s latent space with endoscopy-specific visual features. The method establishes the largest endoscopic generative foundation model to date, enabling high-fidelity image synthesis and efficient editing capabilities such as inpainting and extrapolation. Generated images exhibit rich detail and structural coherence, achieving classification performance on multiple benchmarks that matches or even surpasses that of specialized endoscopic foundation models, thereby significantly accelerating downstream task development.
📝 Abstract
Developing foundation generative models for endoscopy is limited by the gap between natural and clinical images and the computational cost of training large Diffusion Transformers. Although representation alignment has improved efficiency in general computer vision, its role within the highly specialized endoscopic image space remains unclear. We introduce REVEAL (Representation-driven Endoscopic Visual Embedding Alignment), the largest generative foundation model for endoscopy to date, trained on GastroNet-5M (GN-5M), a multicenter dataset of 5 million endoscopic frames. Instead of depending on out-of-domain priors, REVEAL employs encoders pretrained directly on the endoscopic distribution to align diffusion latents with domain-specific visual features, preserving fine textures and intricate anatomical structures. Beyond image generation, REVEAL also serves as a powerful feature extractor; in multiple benchmarks, it delivers performance that is competitive with, and in several cases exceeds, endoscopic foundation models such as EndoViT and Endo-FM, specifically tuned for classification tasks, while demonstrating strong representation robustness under realistic imaging corruptions. REVEAL produces high-fidelity images and maintains robust structural coherence in latent-space edits such as inpainting and outpainting. This high-capacity backbone lowers the computational threshold for building specialized clinical tools, offering an open, versatile foundation for conditional synthesis, segmentation, and out-of-distribution detection in future intelligent gastroenterology systems.
Problem

Research questions and friction points this paper is trying to address.

endoscopy
generative foundation model
representation alignment
diffusion transformer
domain gap
Innovation

Methods, ideas, or system contributions that make the work stand out.

representation alignment
endoscopic foundation model
diffusion latent
domain-specific pretraining
latent-space editing
🔎 Similar Papers
No similar papers found.
F
Francisco Caetano
Department of Electrical Engineering, ARIA Lab, Eindhoven University of Technology, Eindhoven, The Netherlands
T
Tim J. M. Jaspers
Department of Electrical Engineering, ARIA Lab, Eindhoven University of Technology, Eindhoven, The Netherlands
H
Haiko Middeljans
Department of Electrical Engineering, ARIA Lab, Eindhoven University of Technology, Eindhoven, The Netherlands
M
Martijn R. Jong
Department of Gastroenterology and Hepatology, Amsterdam University Medical Centers, University of Amsterdam, Amsterdam, The Netherlands
R
Rixta A. H. van Eijck van Heslinga
Department of Gastroenterology and Hepatology, Amsterdam University Medical Centers, University of Amsterdam, Amsterdam, The Netherlands
F
Floor Slooter
Department of Gastroenterology and Hepatology, Amsterdam University Medical Centers, University of Amsterdam, Amsterdam, The Netherlands
A
Albert J. de Groof
Department of Gastroenterology and Hepatology, Amsterdam University Medical Centers, University of Amsterdam, Amsterdam, The Netherlands
J
Jacques J. Bergman
Department of Gastroenterology and Hepatology, Amsterdam University Medical Centers, University of Amsterdam, Amsterdam, The Netherlands
P
Peter H. N. De With
Department of Electrical Engineering, ARIA Lab, Eindhoven University of Technology, Eindhoven, The Netherlands
Fons van der Sommen
Fons van der Sommen
Associate Professor, Eindhoven University of Technology
Image processingComputer VisionMedical Image AnalysisComputer-Aided DiagnosisMachine learning