Energy Variation in Training Modern Computer Vision Architectures

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the surging energy consumption in deep learning training and the lack of empirical evidence regarding the energy efficiency of modern vision architectures. Utilizing a dual Tesla P100 cluster with NVML monitoring, we systematically quantify the training energy consumption of seven architectures—including MobileNet, EfficientNet, ViT, and ConvNeXt—on ImageNet. Our analysis reveals a strong correlation between GFLOPs and energy consumption (r=0.85), with up to a 3.1-fold variation in energy usage under equivalent computational budgets. Furthermore, we empirically validate the energy inefficiency of Vision Transformers and identify EfficientNet as the optimal accuracy-efficiency trade-off. This work fills a critical gap in the energy efficiency evaluation of computer vision architectures, providing an empirical foundation for model selection in resource-constrained scenarios.
📝 Abstract
The rapid growth of deep learning has substantially increased the energy consumption associated with model training, making energy efficiency an increasingly relevant design criterion. This study empirically measures the energy variation of training seven modern computer vision architectures, MobileNetV3-Small, MobileNetV3-Large, EfficientNet-B0, EfficientNet-B1, ViT-B/32, ConvNeXt-Tiny, and ViT-B/16 for the ImageNet-1k classification task, using a homogeneous 10,000-image subset (ImageNet-10k) and a uniform 40-epoch baseline configuration executed on two NVIDIA Tesla P100 GPUs at the Bioinformatics and Computational Biology Center of Colombia (BIOS). Energy was recorded directly via NVML and contrasted with the computational complexity of each model. The results show a Pearson correlation of 0.85 between floating-point operations (GFLOPs) and energy consumption in kWh, indicating that computational complexity is a strong but imperfect predictor of energy expenditure: architectures with comparable GFLOPs exhibited consumption differing by up to 3.1x due to differences in the hardware efficiency of their dominant operations. The EfficientNet variants offered the best balance between classification performance (Val Top-5 up to 97.15%) and energy efficiency (0.457-0.611 kWh), while Vision Transformers exhibited the highest relative energy consumption and lower classification performance under the evaluated configuration. These findings guide architecture selection in energy-constrained computing environments.
Problem

Research questions and friction points this paper is trying to address.

Energy Consumption
Computer Vision Architectures
Model Training
Computational Complexity
Energy Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Energy Efficiency
Computer Vision Architectures
Computational Complexity
Empirical Measurement
Vision Transformers
🔎 Similar Papers
No similar papers found.