🤖 AI Summary
This study addresses the model selection challenge for near-real-time breast implant detection by systematically evaluating the performance trade-offs between foundation models, including RAD-DINO and MammoCLIP, and specialized ResNet-based CNNs. Using the Emory dataset, we process embedding features via support vector machines with grid-search-optimized architectures and propose a novel lightweight variant, ResNetLite. Results demonstrate that MammoCLIP achieves the highest AUROC (0.999) and the fastest training speed in GPU environments. Notably, ResNetLite attains statistically comparable accuracy using only 1.4% of the parameters while delivering the fastest inference, making it highly suitable for edge deployment. By delineating these architectural trade-offs, this work provides clear model selection strategies tailored to diverse clinical scenarios, balancing computational constraints with diagnostic performance requirements.
📝 Abstract
Purpose: To evaluate the performance-feasibility tradeoffs of foundation models (FMs) and task-specific convolutional neural networks (CNNs) trained from scratch for breast implant classification in 2D mammography, with emphasis on suitability for near real-time clinical deployment. Methods: We evaluated four models: two FMs (RAD-DINO and MammoCLIP) and two CNNs trained from scratch for implant prediction (ResNet18 and our lightweight ResNetLite). Using the Emory Breast Imaging Dataset, 5,000 unilateral screening mammograms were used for training/validation and 1,000 manually reviewed unilateral images were held out for testing. For the FMs, global image embeddings from the pretrained encoder were classified using a support vector machine (SVM). The CNNs were trained end-to-end on 2D mammograms, with ResNetLite optimized via grid search over depth and width to balance accuracy and efficiency. Performance was evaluated using AUROC, sensitivity, specificity, accuracy, embedding visualization, and inference-latency. Results: All models demonstrated strong performance on held-out test data (n = 1,000). MammoCLIP achieved the highest AUROC (0.999) with the quickest training time of 493 seconds. RAD-DINO achieved the highest sensitivity (0.980; accuracy 0.989) but had the slowest inference and training times. ResNet18 and MammoCLIP achieved comparable accuracy (0.985). ResNetLite showed no statistically significant difference from ResNet18 (AUROC 0.993; accuracy 0.976) despite using only 1.4% of ResNet18's parameters, and had the fastest inference time. Conclusion: FMs and task-specific CNN models reliably detect breast implants on 2D mammography. Model selection is best guided by deployment context: MammoCLIP for GPU-equipped hospital settings requiring scalable integration, and lightweight CNNs such as ResNetLite for resource-constrained or edge deployments.