design lightweight convnets

Design, build, or analyze compact convolutional neural network architectures that minimize parameter counts, FLOPs, memory, and latency to enable real-time on-device inference by applying techniques such as depthwise and other separable convolutions, residual blocks, and squeeze-and-excitation or other attention modules. Optimize layer choices, architecture topology, and hyperparameters to preserve or improve accuracy and spatial feature fidelity while keeping computational and memory footprint within tight resource budgets.

designlightweightconvnets

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.46
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work proposes a parameterized convolutional accelerator architecture based on high-level synthesis (HLS) to address the limitations of conventional CNN accelerators, which often prioritize peak performance at the expense of critical embedded constraints such as latency, power consumption, area, and cost. By leveraging a hardware-software co-design approach, the proposed architecture enables efficient multi-objective optimization across these dimensions, overcoming the rigidity of fixed architectures. Experimental results demonstrate that, compared to non-parameterized designs, the proposed solution not only meets stringent embedded deployment requirements but also offers superior scalability and energy efficiency. Furthermore, the framework exhibits broad applicability and can be readily extended to other deep learning acceleration scenarios.

CNN acceleratordesign constraintsembedded deep learning

Exploring Quantization and Mapping Synergy in Hardware-Aware Deep Neural Network Accelerators

Apr 03, 2024
JK
Jan Klhufek
🏛️ Brno University of Technology

This work addresses the disjoint optimization of quantization and hardware mapping in CNN accelerators, where energy efficiency and memory constraints are tightly coupled. We propose a quantization–mapping co-design methodology that jointly optimizes weight quantization policies and hardware mapping (scheduling and resource allocation) under multiple objectives. We extend the Timeloop framework to support mixed-precision quantization modeling and introduce a layer-wise adaptive bit-width–mapping co-search algorithm. Evaluated on Eyeriss and Simba architectures with MobileNetV1/V2, our approach achieves up to 37% energy reduction on ImageNet with zero accuracy loss, significantly expanding the Pareto frontier across energy efficiency, accuracy, and memory usage. Our core contribution is the identification of a novel, high-efficiency mapping space enabled by mixed-precision quantization and the development of the first open-source toolchain supporting quantization-aware, joint mapping optimization.

Develop mixed quantization support in acceleratorsExplore quantization and mapping synergyOptimize CNN energy and memory efficiency

Compact neural networks often suffer from low hardware efficiency due to redundant activation functions and sparse, hardware-unfriendly operators (e.g., depthwise convolutions), leading to poor utilization of compute resources. Method: This paper introduces “structural shrinking,” a novel post-training compression paradigm that identifies and safely removes redundant activation functions; it further combines activation pruning, operator reparameterization, and dense block reconstruction to convert irregular sparse architectures into highly parallel, throughput-optimized dense computation modules. Contribution/Results: The proposed hardware-aware compression framework achieves a 1.53× throughput improvement over MetaPruning on Tesla V100 GPUs while simultaneously improving accuracy by 3.06%, demonstrating superior co-optimization of hardware efficiency and model accuracy.

Computational OptimizationEfficient Neural NetworksHardware Utilization

A flexible FPGA accelerator for convolutional neural networks

Dec 16, 2019
KM
Kingshuk Majumder
🏛️ Indian Institute of Science

To address the memory bandwidth bottleneck and hardware programming complexity in FPGA-accelerated CNN inference, this paper proposes a scalable software-hardware co-design acceleration architecture. Methodologically: (1) it introduces a reconfiguration-free 1D processing element (PE) array with inter-layer adaptive scheduling, eliminating resource waste caused by non-aligned layer dimensions; (2) it establishes a multi-level data reuse mechanism and a reconfiguration-free core array to maximize on-chip memory and computational resource utilization; (3) it presents the first TensorFlow-front-end-supported software-hardware co-compilation framework, enabling automatic tiling, dataflow scheduling, and RTL generation. Evaluated on the Xilinx VC709 platform, the architecture achieves a high sustained performance ratio relative to its theoretical peak, with significant improvements in throughput and energy efficiency. It further demonstrates high-frequency scalability and strong engineering practicality.

Accelerating CNN inference on FPGAs efficientlyMinimizing off-chip memory access through reuse optimizationOvercoming FPGA programming difficulty via TensorFlow integration

An Efficient Real-Time Object Detection Framework on Resource-Constricted Hardware Devices via Software and Hardware Co-design

Jul 01, 2021
SL
Shiyi Luo
🏛️ San Diego State University | University of California Irvine | University of California, Davis | California State University Fullerton

Existing tensor decomposition-based rank selection for embedded devices relies heavily on manual trial-and-error or incurs prohibitive computational overhead from automatic optimization. To address this, we propose a software-hardware co-designed real-time object detection framework. Our approach uniquely integrates Tensor Train (TT) decomposition with FPGA acceleration in a deeply coupled manner, enabling joint optimization of model compression ratio and hardware execution efficiency. Specifically, we apply TT decomposition to compress YOLOv5, design a custom FPGA accelerator, and perform software-hardware co-compiled optimizations. Evaluated on Jetson Nano and Xilinx Zynq FPGA platforms, the framework achieves 68% model size reduction, 3.2× inference speedup, and end-to-end latency under 32 ms—while preserving high detection accuracy. This work establishes a scalable, co-design paradigm for efficient, lightweight vision models at the edge.

Automating tensor rank selection for neural network compressionBalancing compression efficiency with minimal accuracy lossReducing computational complexity in model compression methods

Latest Papers

What's happening recently
View more

This work addresses the challenge of deploying convolutional neural networks (CNNs) on memory-constrained microcontroller units (MCUs), where high peak RAM usage—primarily due to intermediate activation tensors during inference—prevents standalone execution. To overcome this limitation, the authors propose a fine-grained collaborative inference system that moves beyond conventional layer-wise model partitioning by decomposing networks at the granularity of individual convolutional kernels or neurons. A lightweight, resource-aware coordinator dynamically schedules computations across heterogeneous MCUs, enabling efficient utilization of distributed resources. The approach successfully deploys previously infeasible models such as MobileNetV2 on platforms comprising up to eight MCUs, substantially reducing peak memory consumption per MCU while maintaining practical end-to-end inference latency.

CNN inferenceintermediate activationsmemory bottleneck

Deploying deep neural networks on edge devices entails balancing accuracy, latency, and resource constraints. This work presents the first end-to-end hardware evaluation comparing static compression techniques—namely pruning and quantization—with dynamic early-exit mechanisms, all implemented within a unified ONNX inference framework. Experimental results demonstrate that static methods substantially reduce memory footprint, while early-exit strategies achieve input-adaptive computational savings. Crucially, combining both approaches yields simultaneous reductions in both latency and memory consumption with negligible accuracy loss, revealing their complementary nature and significant potential for joint optimization in edge computing scenarios.

DeploymentEarly ExitEdge AI

This work addresses the limitations of existing hardware-aware neural architecture search (HW-NAS) methods, which are primarily tailored for high-performance microcontrollers and fail to meet the stringent resource constraints of ultra-low-power sensing nodes. To bridge this gap, the authors propose a lightweight HW-NAS framework specifically designed for ultra-low-power microcontrollers, enabling, for the first time, fully on-device, end-to-end searchable deployment of tiny convolutional neural networks directly on embedded platforms. By jointly optimizing model accuracy and hardware-specific constraints, the method demonstrates consistent effectiveness across three mainstream micro-vision benchmarks. The resulting architectures achieve state-of-the-art classification accuracy while being successfully deployable on ultra-low-power hardware, thereby advancing the feasibility of intelligent edge inference under extreme energy budgets.

convolutional neural networksembedded deviceshardware-aware neural architecture search

This work addresses the challenge of performing end-to-end on-device training and inference for vision-based machine learning on extremely resource-constrained microcontrollers, where traditional approaches rely heavily on cloud infrastructure. The authors demonstrate a complete pipeline—including data acquisition, training of a two-layer convolutional neural network (CNN), and real-time inference—on an ESP32-S3 XIAO ML Kit featuring only 8 MB of PSRAM and no external dependencies. By leveraging batch-level gradient accumulation, precomputed scaling lookup tables, a three-priority weight loading scheme, and PSRAM-aware memory management, the system achieves a full training cycle in just 9 minutes and inference at 6.3 frames per second on a 64×64 three-class classification task. Implemented in only 1,750 lines of C++ code, the system includes a custom Adam optimizer and CNN compatible with the Arduino IDE, and is released under the MIT license.

Edge ComputingMicrocontrollerOn-Device Learning

Hot Scholars

RT

Radu Timofte

Humboldt Professor for AI and Computer Vision, University of Würzburg
Computer VisionMachine LearningAICompression
ZW

Zongwei Wu

University of Würzburg | CNRS - Université de Bourgogne | ETH Zurich
Sensor FusionPerception
QH

Qibin Hou

Nankai University
Deep learningComputer visionVisual attention
MM

Ming-Ming Cheng

Professor of Computer Science, Nankai University
Computer VisionComputer GraphicsVisual AttentionSaliency
JW

Jun-Wei Hsieh

National Yang Ming Chiao Tung University
computer visionAIimage processing