CNN and ViT Efficiency Study on Tiny ImageNet and DermaMNIST Datasets

πŸ“… 2025-05-13
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study investigates the accuracy-efficiency trade-off between CNNs and Vision Transformers (ViTs) on medical (DermaMNIST) and general-purpose (TinyImageNet) image classification tasks. Using ResNet-18 as a baseline, we systematically fine-tune four ViT variants under strict constraints of low latency and parameter count. Methodologically, we introduce a lightweight fine-tuning protocol tailored for resource-constrained deployment. Our key contribution is the first empirical demonstration that minimally adapted ViTs can outperform CNN baselines on cross-domain, small-scale medical data. Results show that ViT-Small achieves a 1.2% accuracy gain on DermaMNIST with 37% fewer parameters and 1.8Γ— faster inference; ViT-Tiny attains 99.4% of ResNet-18’s TinyImageNet accuracy while using only 42% of its parameters. The work establishes ViTs’ viability in low-data medical settings and provides a transferable lightweight adaptation framework for efficient vision model deployment.

Technology Category

Computer Vision: Large Vision ModelsMachine Learning: Transfer, Domain Adaptation, Multi-Task LearningNatural Language Processing: Safety and Robustness

Application Category

Web Mining and Content Analysis: Large pretrained models with web dataSearch and Retrieval-Augmented AI: Efficiency and scalability of Web search enginesSystems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applications
πŸ“ Abstract
This study evaluates the trade-offs between convolutional and transformer-based architectures on both medical and general-purpose image classification benchmarks. We use ResNet-18 as our baseline and introduce a fine-tuning strategy applied to four Vision Transformer variants (Tiny, Small, Base, Large) on DermatologyMNIST and TinyImageNet. Our goal is to reduce inference latency and model complexity with acceptable accuracy degradation. Through systematic hyperparameter variations, we demonstrate that appropriately fine-tuned Vision Transformers can match or exceed the baseline's performance, achieve faster inference, and operate with fewer parameters, highlighting their viability for deployment in resource-constrained environments.
Problem

Research questions and friction points this paper is trying to address.

Compare CNN and ViT efficiency on medical and general image datasets
Reduce inference latency and model complexity with minimal accuracy loss
Evaluate fine-tuned ViT variants for resource-constrained deployment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fine-tuned Vision Transformers for efficient image classification
Reduced inference latency with acceptable accuracy trade-offs
Optimized model complexity for resource-constrained environments
πŸ”Ž Similar Papers
2024-08-29Medical Imaging 2025: Digital and Computational PathologyCitations: 1
πŸ’Ό Related Jobs
No related jobs found.
A
Aidar Amangeldi
Department of Data Science, Nazarbayev University, Astana, Kazakhstan
A
Angsar Taigonyrov
Department of Computer Science, Nazarbayev University, Astana, Kazakhstan
M
Muhammad Huzaid Jawad
Department of Data Science, Nazarbayev University, Astana, Kazakhstan
C
Chinedu Emmanuel Mbonu
Department of Computer Science, Nazarbayev University, Astana, Kazakhstan