π€ AI Summary
This study investigates the accuracy-efficiency trade-off between CNNs and Vision Transformers (ViTs) on medical (DermaMNIST) and general-purpose (TinyImageNet) image classification tasks. Using ResNet-18 as a baseline, we systematically fine-tune four ViT variants under strict constraints of low latency and parameter count. Methodologically, we introduce a lightweight fine-tuning protocol tailored for resource-constrained deployment. Our key contribution is the first empirical demonstration that minimally adapted ViTs can outperform CNN baselines on cross-domain, small-scale medical data. Results show that ViT-Small achieves a 1.2% accuracy gain on DermaMNIST with 37% fewer parameters and 1.8Γ faster inference; ViT-Tiny attains 99.4% of ResNet-18βs TinyImageNet accuracy while using only 42% of its parameters. The work establishes ViTsβ viability in low-data medical settings and provides a transferable lightweight adaptation framework for efficient vision model deployment.
π Abstract
This study evaluates the trade-offs between convolutional and transformer-based architectures on both medical and general-purpose image classification benchmarks. We use ResNet-18 as our baseline and introduce a fine-tuning strategy applied to four Vision Transformer variants (Tiny, Small, Base, Large) on DermatologyMNIST and TinyImageNet. Our goal is to reduce inference latency and model complexity with acceptable accuracy degradation. Through systematic hyperparameter variations, we demonstrate that appropriately fine-tuned Vision Transformers can match or exceed the baseline's performance, achieve faster inference, and operate with fewer parameters, highlighting their viability for deployment in resource-constrained environments.