🤖 AI Summary
This study addresses the scarcity of annotated data for ultrasound-based pulmonary tuberculosis screening by investigating whether domain-specific pretraining on cardiac ultrasound, which shares underlying physical principles, outperforms general video pretraining. Employing deep learning and transfer learning frameworks, we systematically compare multiple encoders, including V-JEPA and EchoJEPA, on lung ultrasound tasks, with particular emphasis on the role of feature standardization. Our findings reveal that shared physical priors do not necessarily enhance transferability; rather, feature standardization emerges as the critical factor driving performance improvements. Experimental results demonstrate that standardized encoders achieve an average performance gain of 1.23%, with a maximum AUROC improvement of 2.57% and a specificity reaching 79.3%. This work establishes feature standardization as a core mechanism for optimizing encoder performance in medical ultrasound transfer learning scenarios.
📝 Abstract
Lung ultrasound (LUS) is attractive for tuberculosis (TB) screening at primary-care level, but labelled cohorts are small. Echocardiography carries no such constraint, while sharing the same underlying ultrasound imaging physics, signal processing and B-mode appearance as LUS. We ask whether an encoder pretrained on that high-resource ultrasound domain carries representations that remain usable in the low-resource one. Only the encoder varies, across seventeen encoders spanning three architecture families. Among them, a latent-predictive video encoder pretrained on generic video (V-JEPA2-L) and its echocardiography counterpart (EchoJEPA-L) differ in pretraining corpus alone. The choice among these encoders does not resolve the classification, the whole family spanning 2.50 percentage points against a measurement resolution of 2.71. What moves the task instead is feature conditioning. Standardising the features between the encoder and the classifier improves all seventeen encoders by a mean of +1.23 percentage points at $p=1.5\times10^{-5}$. On the held-out test set every encoder selected on the development folds stands above the baseline system by up to +2.57 percentage points of area under the receiver operating characteristic curve (AUROC), and specificity at 90% sensitivity reaches 79.3% against 60.3%. The contrast specified in advance, EchoJEPA-L against V-JEPA2-L, measures -0.16 percentage points at $p=0.926$. We therefore find no evidence that shared ultrasonic physics alone makes echocardiography a more productive pretraining corpus than generic video, and any advantage, if present, is smaller than this cohort can resolve. The video encoders receive replicated still images, however, so whether this absence of an effect reflects the pretraining domain or a video encoder applied to static frames cannot be separated. The limiting factor is the labelled cohort rather than the encoder.