Score
Design, build, or analyze compact convolutional neural network architectures that minimize parameter counts, FLOPs, memory, and latency to enable real-time on-device inference by applying techniques such as depthwise and other separable convolutions, residual blocks, and squeeze-and-excitation or other attention modules. Optimize layer choices, architecture topology, and hyperparameters to preserve or improve accuracy and spatial feature fidelity while keeping computational and memory footprint within tight resource budgets.
Deep convolutional neural networks (CNNs) face persistent challenges in balancing representational capacity with deployment efficiency across diverse domains—including vision, language, healthcare, and speech—especially under resource constraints and data-limited regimes. Method: This work systematically surveys CNN architectural evolution from 2015 to 2025 and proposes a novel seven-dimensional unified taxonomy (covering spatial modeling, multi-path design, dimensional expansion, attention integration, etc.), alongside a synergistic optimization framework integrating sparse convolutions, depthwise separable convolutions, and attention mechanisms. It further incorporates Fourier-based preprocessing, low-precision computation, and weight compression for lightweight deployment. Contribution/Results: We comprehensively characterize the applicability boundaries of over 100 CNN variants, quantify the trade-off between computational efficiency and representation fidelity, formalize adaptation strategies for few-shot, weakly supervised, and federated learning settings, and prospectively identify emerging directions—including CNN-Transformer hybrids, vision-language joint modeling, and generative CNNs—establishing reusable design paradigms for edge deployment and cross-domain generalization.
This work proposes a parameterized convolutional accelerator architecture based on high-level synthesis (HLS) to address the limitations of conventional CNN accelerators, which often prioritize peak performance at the expense of critical embedded constraints such as latency, power consumption, area, and cost. By leveraging a hardware-software co-design approach, the proposed architecture enables efficient multi-objective optimization across these dimensions, overcoming the rigidity of fixed architectures. Experimental results demonstrate that, compared to non-parameterized designs, the proposed solution not only meets stringent embedded deployment requirements but also offers superior scalability and energy efficiency. Furthermore, the framework exhibits broad applicability and can be readily extended to other deep learning acceleration scenarios.
This work addresses the disjoint optimization of quantization and hardware mapping in CNN accelerators, where energy efficiency and memory constraints are tightly coupled. We propose a quantization–mapping co-design methodology that jointly optimizes weight quantization policies and hardware mapping (scheduling and resource allocation) under multiple objectives. We extend the Timeloop framework to support mixed-precision quantization modeling and introduce a layer-wise adaptive bit-width–mapping co-search algorithm. Evaluated on Eyeriss and Simba architectures with MobileNetV1/V2, our approach achieves up to 37% energy reduction on ImageNet with zero accuracy loss, significantly expanding the Pareto frontier across energy efficiency, accuracy, and memory usage. Our core contribution is the identification of a novel, high-efficiency mapping space enabled by mixed-precision quantization and the development of the first open-source toolchain supporting quantization-aware, joint mapping optimization.
Compact neural networks often suffer from low hardware efficiency due to redundant activation functions and sparse, hardware-unfriendly operators (e.g., depthwise convolutions), leading to poor utilization of compute resources. Method: This paper introduces “structural shrinking,” a novel post-training compression paradigm that identifies and safely removes redundant activation functions; it further combines activation pruning, operator reparameterization, and dense block reconstruction to convert irregular sparse architectures into highly parallel, throughput-optimized dense computation modules. Contribution/Results: The proposed hardware-aware compression framework achieves a 1.53× throughput improvement over MetaPruning on Tesla V100 GPUs while simultaneously improving accuracy by 3.06%, demonstrating superior co-optimization of hardware efficiency and model accuracy.
To address the memory bandwidth bottleneck and hardware programming complexity in FPGA-accelerated CNN inference, this paper proposes a scalable software-hardware co-design acceleration architecture. Methodologically: (1) it introduces a reconfiguration-free 1D processing element (PE) array with inter-layer adaptive scheduling, eliminating resource waste caused by non-aligned layer dimensions; (2) it establishes a multi-level data reuse mechanism and a reconfiguration-free core array to maximize on-chip memory and computational resource utilization; (3) it presents the first TensorFlow-front-end-supported software-hardware co-compilation framework, enabling automatic tiling, dataflow scheduling, and RTL generation. Evaluated on the Xilinx VC709 platform, the architecture achieves a high sustained performance ratio relative to its theoretical peak, with significant improvements in throughput and energy efficiency. It further demonstrates high-frequency scalability and strong engineering practicality.
Existing tensor decomposition-based rank selection for embedded devices relies heavily on manual trial-and-error or incurs prohibitive computational overhead from automatic optimization. To address this, we propose a software-hardware co-designed real-time object detection framework. Our approach uniquely integrates Tensor Train (TT) decomposition with FPGA acceleration in a deeply coupled manner, enabling joint optimization of model compression ratio and hardware execution efficiency. Specifically, we apply TT decomposition to compress YOLOv5, design a custom FPGA accelerator, and perform software-hardware co-compiled optimizations. Evaluated on Jetson Nano and Xilinx Zynq FPGA platforms, the framework achieves 68% model size reduction, 3.2× inference speedup, and end-to-end latency under 32 ms—while preserving high detection accuracy. This work establishes a scalable, co-design paradigm for efficient, lightweight vision models at the edge.
This work addresses the challenge of deploying convolutional neural networks (CNNs) on memory-constrained microcontroller units (MCUs), where high peak RAM usage—primarily due to intermediate activation tensors during inference—prevents standalone execution. To overcome this limitation, the authors propose a fine-grained collaborative inference system that moves beyond conventional layer-wise model partitioning by decomposing networks at the granularity of individual convolutional kernels or neurons. A lightweight, resource-aware coordinator dynamically schedules computations across heterogeneous MCUs, enabling efficient utilization of distributed resources. The approach successfully deploys previously infeasible models such as MobileNetV2 on platforms comprising up to eight MCUs, substantially reducing peak memory consumption per MCU while maintaining practical end-to-end inference latency.
Deploying deep neural networks on edge devices entails balancing accuracy, latency, and resource constraints. This work presents the first end-to-end hardware evaluation comparing static compression techniques—namely pruning and quantization—with dynamic early-exit mechanisms, all implemented within a unified ONNX inference framework. Experimental results demonstrate that static methods substantially reduce memory footprint, while early-exit strategies achieve input-adaptive computational savings. Crucially, combining both approaches yields simultaneous reductions in both latency and memory consumption with negligible accuracy loss, revealing their complementary nature and significant potential for joint optimization in edge computing scenarios.
This work addresses the limitations of existing hardware-aware neural architecture search (HW-NAS) methods, which are primarily tailored for high-performance microcontrollers and fail to meet the stringent resource constraints of ultra-low-power sensing nodes. To bridge this gap, the authors propose a lightweight HW-NAS framework specifically designed for ultra-low-power microcontrollers, enabling, for the first time, fully on-device, end-to-end searchable deployment of tiny convolutional neural networks directly on embedded platforms. By jointly optimizing model accuracy and hardware-specific constraints, the method demonstrates consistent effectiveness across three mainstream micro-vision benchmarks. The resulting architectures achieve state-of-the-art classification accuracy while being successfully deployable on ultra-low-power hardware, thereby advancing the feasibility of intelligent edge inference under extreme energy budgets.
This work addresses the challenge of performing end-to-end on-device training and inference for vision-based machine learning on extremely resource-constrained microcontrollers, where traditional approaches rely heavily on cloud infrastructure. The authors demonstrate a complete pipeline—including data acquisition, training of a two-layer convolutional neural network (CNN), and real-time inference—on an ESP32-S3 XIAO ML Kit featuring only 8 MB of PSRAM and no external dependencies. By leveraging batch-level gradient accumulation, precomputed scaling lookup tables, a three-priority weight loading scheme, and PSRAM-aware memory management, the system achieves a full training cycle in just 9 minutes and inference at 6.3 frames per second on a 64×64 three-class classification task. Implemented in only 1,750 lines of C++ code, the system includes a custom Adam optimizer and CNN compatible with the Arduino IDE, and is released under the MIT license.