Score
Designs and builds convolutional neural network architectures that extract features at multiple spatial scales by combining parallel convolutional branches (e.g., 1×1, 3×3, 5×5) and Inception-style modules, while producing end-to-end trainable models. This includes specifying layer types, connectivity and bottlenecks, and applying parameter- and compute-efficient choices (e.g., global average pooling, dimensionality reduction) to balance capacity for texture/edge/shape representation with resource constraints.
Deep convolutional neural networks (CNNs) face persistent challenges in balancing representational capacity with deployment efficiency across diverse domains—including vision, language, healthcare, and speech—especially under resource constraints and data-limited regimes. Method: This work systematically surveys CNN architectural evolution from 2015 to 2025 and proposes a novel seven-dimensional unified taxonomy (covering spatial modeling, multi-path design, dimensional expansion, attention integration, etc.), alongside a synergistic optimization framework integrating sparse convolutions, depthwise separable convolutions, and attention mechanisms. It further incorporates Fourier-based preprocessing, low-precision computation, and weight compression for lightweight deployment. Contribution/Results: We comprehensively characterize the applicability boundaries of over 100 CNN variants, quantify the trade-off between computational efficiency and representation fidelity, formalize adaptation strategies for few-shot, weakly supervised, and federated learning settings, and prospectively identify emerging directions—including CNN-Transformer hybrids, vision-language joint modeling, and generative CNNs—establishing reusable design paradigms for edge deployment and cross-domain generalization.
To address excessive computational overhead when deploying vision models on resource-constrained devices, this paper proposes an efficient Vision Transformer (ViT) architecture design framework. Methodologically: (1) it optimizes the input-output data pathway to enhance representational capacity of lightweight models; (2) it restructures the context window of computationally constrained attention mechanisms to improve local-global modeling efficiency; and (3) it leverages the invertibility and explicit probabilistic modeling properties of normalizing flows to enable high-fidelity, low-overhead knowledge distillation. Experiments demonstrate that the proposed approach achieves comparable or superior accuracy on benchmarks such as ImageNet, while requiring significantly fewer parameters and FLOPs. It also substantially reduces inference latency and memory footprint. The framework establishes a scalable new paradigm for efficient visual understanding at the edge.
Tropical convolutional neural networks (TCNNs) suffer from low accuracy, whereas standard CNNs incur high computational overhead. Method: This paper proposes two novel tropical convolution architectures—composite (cTCNN) and parallel (pTCNN)—replacing conventional convolutional kernels with min-plus/max-plus kernel compositions. It introduces, for the first time, composite and parallel tropical convolution patterns into deep networks and constructs a hybrid architecture synergistically integrating tropical algebra with standard CNNs. Contribution/Results: The approach achieves near-zero multiplicative complexity while significantly improving accuracy. Experiments on multiple benchmark datasets show that the proposed models match or surpass state-of-the-art CNNs in accuracy; cascading them with standard CNNs further enhances deep model performance; and they reduce parameter count and multiplication operations substantially—with negligible accuracy degradation—achieving state-of-the-art lightweight efficiency.
This study investigates the selection of efficient convolutional neural network (CNN) architectures for image classification and object detection under varying task complexity and resource constraints. Through systematic comparisons across five real-world datasets—including binary classification, fine-grained multiclass classification, and object detection tasks—the work evaluates custom CNNs, deep residual networks, and transfer learning models, analyzing the impact of key architectural factors such as network depth and residual connections. The results demonstrate that deeper architectures significantly improve accuracy in fine-grained classification, whereas lightweight pretrained models offer superior efficiency for simpler binary classification tasks. Furthermore, the proposed custom CNN is successfully extended to detect illegally operating tricycles in traffic scenarios, confirming its practical effectiveness in real-world applications.
This work investigates the fundamental differences between spatial token mixers (STMs)—the spatial feature aggregation mechanisms—in Vision Transformers and convolutional networks. To enable a fair, architecture-agnostic comparison, we propose a unified STM modeling paradigm that decouples network-level design from the spatial aggregation module, implementing both convolutional and attention-based STMs on a neutral backbone. Our methodology includes: (1) designing a modular, swappable STM interface; (2) systematically analyzing inductive biases—including receptive field size, translation invariance, and adversarial robustness; and (3) conducting multi-task performance benchmarking. Results show that while modern network-level designs yield substantial gains, intrinsic performance gaps among STMs persist. Crucially, we quantitatively demonstrate for the first time that convolutions exhibit superior translation invariance and local robustness, whereas attention achieves larger effective receptive fields but is more vulnerable to input perturbations.
Traditional hand-crafted histogram features—such as Local Binary Patterns (LBP) and edge histograms—are incompatible with end-to-end deep learning due to their non-differentiability. To address this, we propose a differentiable histogram layer, enabling the first neuralization and learnability of such features. Methodologically, we design Neural Local Binary Patterns (NLBP) and Neural Edge Histogram Descriptor (NEHD) modules, integrated as differentiable statistical layers within CNNs to support gradient backpropagation and joint optimization. Our core contribution lies in unifying hand-engineered feature design with deep learning paradigms, allowing local statistical priors to be data-drivenly learned and enhanced. Extensive experiments on multiple image classification benchmarks and real-world datasets demonstrate consistent and significant performance gains, validating that neuralized histogram features substantially improve representation capability.
This work proposes a lightweight, customized CNN architecture to address the significant disparities between agricultural and urban scene images—particularly in illumination, resolution, environmental complexity, and class imbalance—and to build an efficient, robust general-purpose visual classification model. Through systematic comparisons with mainstream architectures such as ResNet-18 and VGG-16 across five heterogeneous datasets, the study evaluates the proposed model’s convergence behavior, generalization capability, and performance under both from-scratch training and transfer learning settings across varying data scales. Experimental results demonstrate that the custom CNN achieves accuracy comparable to established models while maintaining a compact footprint. Furthermore, this study provides the first systematic characterization of the practical performance boundaries of transfer learning in small-sample, highly heterogeneous scenarios, offering both theoretical insights and practical guidance for deployment under resource constraints.
This work addresses the challenges of high inference latency, substantial memory consumption, and excessive power usage in convolutional neural networks (CNNs) due to their large scale, as well as the lack of a unified framework and comparability among existing pruning methods. To this end, the authors propose Bonsai, a novel pruning framework that establishes a general-purpose platform called Combine, introduces a standardized language for describing pruning criteria, and incorporates several new filter-level pruning strategies. Through an iterative structured pruning approach, Bonsai removes up to 79% of filters in VGG-style models, reduces computational cost by 68%, and maintains or even improves model accuracy. The study systematically elucidates the performance disparities arising from different pruning criteria.
This study investigates the impact of convolutional neural network (CNN) architecture design on image classification performance across agricultural and urban domains. To this end, the authors propose a customized CNN—CustomCNN—that integrates residual connections, Squeeze-and-Excitation attention mechanisms, and a progressive channel scaling strategy, along with Kaiming initialization to enhance representational capacity and training efficiency. Experimental results on five publicly available datasets demonstrate that the proposed model achieves classification performance comparable to that of state-of-the-art CNNs while maintaining computational efficiency. These findings underscore the effectiveness and potential of domain-informed architectural design for visual applications in smart cities and precision agriculture.
Convolutional neural networks (CNNs) suffer from high computational complexity and energy consumption due to dense multiplications, hindering their deployment on resource-constrained mobile devices. To address this, we propose a differentiable lookup table (DLT) operation that replaces multiplication with efficient table lookups while preserving model accuracy. Our key contribution is the first end-to-end differentiable lookup table architecture, seamlessly integrated into CNN layers and jointly optimizable with network parameters via backpropagation. Extensive experiments demonstrate that DLT achieves state-of-the-art performance on image classification, single-image super-resolution, and point cloud classification tasks. It accelerates inference by 1.8–3.2×, reduces energy consumption by 47%–63%, and incurs negligible accuracy degradation (<0.5%). This work establishes a novel paradigm for lightweight CNN design, enabling highly efficient yet accurate deep learning inference on edge devices.