🤖 AI Summary
This study addresses the performance instability of video capsule endoscopy models when deployed across real-world settings involving different devices, annotation protocols, and patient populations due to domain shift. The authors fine-tune general-purpose pretrained backbones on Kvasir-Capsule and systematically evaluate cross-domain generalization using a standardized fine-tuning protocol and decision-space alignment over the shared label subset of the CV2024 and Galar datasets. Their analysis reveals, for the first time, significant inconsistency in model selection across target domains: the model performing best on the source domain exhibits highly variable performance on target domains, with weak inter-domain ranking correlation. The work advocates prioritizing cross-domain ranking stability over peak performance on a single dataset as a more reliable criterion for clinical model selection.
📝 Abstract
Video capsule endoscopy (VCE) classification is typically evaluated within a single dataset, yet clinical deployment demands robustness across acquisition sources, labeling policies, and patient populations. We examine this gap using Kvasir-Capsule, Capsule Vision 2024 (CV2024), and a shared-label subset of Galar. We fine-tune a suite of general-domain pretrained backbones on the official Kvasir-Capsule folds under a standardized protocol and evaluate the same checkpoints on two non-source targets within a documented shared-label decision space. We find that the predictive value of in-domain ranking is target-dependent: Kvasir-Capsule ranking aligns more closely with Galar than with CV2024, while the two non-source targets agree only weakly. Consequently, the strongest in-domain backbone leads on one target yet falls to mid-pack on the other, and no single evaluation target reliably predicts the others. A second CV2024-trained configuration set reproduces this target-dependent instability. We conclude that capsule endoscopy model selection should report cross-target ranking stability rather than peak single-dataset performance.