Score
Designs and trains compact one-dimensional convolutional neural networks that operate directly on raw temporal signals, learning efficient temporal convolutional filters with a small parameter count. Builds, optimizes, and evaluates these models on fixed-length signal windows for tasks such as classification or detection while minimizing computational and memory footprint.
Deep convolutional neural networks (CNNs) face persistent challenges in balancing representational capacity with deployment efficiency across diverse domains—including vision, language, healthcare, and speech—especially under resource constraints and data-limited regimes. Method: This work systematically surveys CNN architectural evolution from 2015 to 2025 and proposes a novel seven-dimensional unified taxonomy (covering spatial modeling, multi-path design, dimensional expansion, attention integration, etc.), alongside a synergistic optimization framework integrating sparse convolutions, depthwise separable convolutions, and attention mechanisms. It further incorporates Fourier-based preprocessing, low-precision computation, and weight compression for lightweight deployment. Contribution/Results: We comprehensively characterize the applicability boundaries of over 100 CNN variants, quantify the trade-off between computational efficiency and representation fidelity, formalize adaptation strategies for few-shot, weakly supervised, and federated learning settings, and prospectively identify emerging directions—including CNN-Transformer hybrids, vision-language joint modeling, and generative CNNs—establishing reusable design paradigms for edge deployment and cross-domain generalization.
Deep learning models for Time Series Classification (TSC) have achieved strong predictive performance but their high computational and memory requirements often limit deployment on resource-constrained devices. While structured pruning can address these issues by removing redundant filters, existing methods typically rely on manually tuned hyperparameters such as pruning ratios which limit scalability and generalization across datasets. In this work, we propose Dynamic Structured Pruning (DSP), a fully automatic, structured pruning framework for convolution-based TSC models. DSP introduces an instance-wise sparsity loss during training to induce channel-level sparsity, followed by a global activation analysis to identify and prune redundant filters without needing any predefined pruning ratio. This work tackles computational bottlenecks of deep TSC models for deployment on resource-constrained devices. We validate DSP on 128 UCR datasets using two different deep state-of-the-art architectures: LITETime and InceptionTime. Our approach achieves an average compression of 58% for LITETime and 75% for InceptionTime architectures while maintaining classification accuracy. Redundancy analyses confirm that DSP produces compact and informative representations, offering a practical path for scalable and efficient deep TSC deployment.
This study addresses the challenge of deploying time-series models on resource-constrained microcontrollers, where computational and memory overheads are significant bottlenecks. Through a hardware-aware comparative analysis, the authors systematically evaluate the end-to-end deployment performance of LSTMs and 1D-CNNs across five time-series classification tasks. Their findings demonstrate, for the first time, that 1D-CNNs consistently outperform LSTMs in TinyML scenarios, achieving an average accuracy of 95%—approximately 6% higher than LSTMs—while simultaneously reducing RAM usage by 35% and Flash consumption by 25%. Moreover, inference latency is dramatically lowered from 2038 ms to 27.6 ms, representing a nearly 74-fold speedup. This work establishes 1D-CNNs as a high-accuracy, low-overhead paradigm for edge-based time-series modeling.
Existing 1D-CNN time-series classification methods for resource-constrained edge devices (e.g., Arduino) require buffering the entire input sequence, resulting in high latency and memory overhead—hindering real-time data acquisition. To address this, we propose a sampling-computation interleaved real-time 1D-CNN inference scheduling mechanism: convolutional operations are dynamically executed within sensor sampling intervals, leveraging a zero-copy circular buffer and a lightweight embedded C-based scheduler to enable streaming inference without redundant data movement. The approach supports cross-platform deployment on both AVR and ARM microcontrollers. On Arduino hardware, it achieves a 10% reduction in end-to-end latency and a 49% decrease in peak memory usage compared to TensorFlow Lite Micro. To our knowledge, this is the first work to realize high-temporal-fidelity, low-overhead online time-series classification on ultra-constrained embedded platforms.
Existing tensor decomposition-based rank selection for embedded devices relies heavily on manual trial-and-error or incurs prohibitive computational overhead from automatic optimization. To address this, we propose a software-hardware co-designed real-time object detection framework. Our approach uniquely integrates Tensor Train (TT) decomposition with FPGA acceleration in a deeply coupled manner, enabling joint optimization of model compression ratio and hardware execution efficiency. Specifically, we apply TT decomposition to compress YOLOv5, design a custom FPGA accelerator, and perform software-hardware co-compiled optimizations. Evaluated on Jetson Nano and Xilinx Zynq FPGA platforms, the framework achieves 68% model size reduction, 3.2× inference speedup, and end-to-end latency under 32 ms—while preserving high detection accuracy. This work establishes a scalable, co-design paradigm for efficient, lightweight vision models at the edge.
This paper investigates the approximation capacity and statistical convergence rates of convolutional neural networks (CNNs) with one-sided zero padding and multi-channel architectures in nonparametric regression and binary classification. Methodologically, it establishes the first compact approximation bound under weight constraints and develops a novel covering-number analysis framework that jointly accounts for network architecture and weight magnitudes, yielding a new upper bound on covering numbers. Consequently, it provides the first theoretical guarantee that the CNN least-squares estimator achieves the minimax optimal rate over Sobolev classes of smooth functions; similarly, CNN classifiers trained with hinge or logistic loss attain minimax optimal rates in binary classification. These results characterize the fundamental statistical efficiency limits of structured CNNs in nonparametric learning, offering rigorous theoretical foundations and practical guidance for architectural design.
This work proposes a learnable inter-filter connectivity mechanism that replaces the fixed pointwise nonlinear activations in conventional convolutional neural networks with a parameterized, universal connection function embedded within convolutional layers. By enabling adaptive interactions among filters, the approach overcomes the limitations of traditional fixed logical operations—such as multiplication or minimum selection—and allows the network to automatically optimize its connectivity strategy through end-to-end training. Experimental results demonstrate that this method significantly improves classification accuracy, confirming its effectiveness in enhancing both model expressivity and generalization capability.
This study addresses the growing concern of first-person-view (FPV) drone misuse in complex electromagnetic environments by proposing a lightweight and efficient radio-frequency signal detection method. The approach leverages software-defined radio to capture drone video transmission signals and directly converts them into time-domain raster images, bypassing conventional spectrogram generation and frequency-domain preprocessing. A compact convolutional neural network is designed to enable end-to-end detection. Experimental evaluation on a dataset of approximately 40,000 annotated images demonstrates that the model achieves high accuracy while significantly reducing computational overhead, exhibiting low latency and a small footprint. The system has been successfully deployed in a real-time monitoring platform on an embedded electronic warfare system.
Traditional CNN training relies on random mini-batch sampling, which often leads to rapid saturation of learning signals as most samples quickly become “easy,” thereby slowing convergence. This work proposes A*-inspired Batch Selection (A*-BS), the first approach to integrate A* search into batch selection, introducing a dynamic scoring mechanism that jointly considers sample loss difficulty and reuse penalties to adaptively select informative and diverse batches. Without modifying network architecture or optimizer, A*-BS achieves superior performance on lightweight CNNs across half of the MedMNIST-v2 benchmark tasks, outperforming ResNet-18/50 in both accuracy and AUC—by up to 15% relatively—while significantly accelerating training. These results demonstrate that intelligent batch sequencing can partially substitute for model depth.
This work addresses the limitations of conventional convolutional time series models in effectively capturing multi-scale structures within long sequences and the tendency of pooling operations to discard critical temporal positional information. The authors propose the ROMAN operator, which explicitly encodes temporal scale and coarse-grained time positions into the channel dimension for the first time. By integrating anti-aliased multi-scale pyramid decomposition, fixed-window slicing, and channel stacking, ROMAN significantly reduces sequence length while preserving multi-scale interactions and temporal positional cues. This design introduces a controllable inductive bias into downstream standard convolutional classifiers such as MiniRocket and FCN. Experiments demonstrate that ROMAN successfully captures coarse-grained positional and multi-scale dependencies on synthetic tasks and substantially improves computational efficiency on long-sequence UCR/UEA benchmarks, with accuracy gains varying across tasks.
This work addresses the high computational cost of forward propagation during the training of deep convolutional neural networks, which significantly limits training efficiency. The authors propose a dynamic layer pruning method tailored specifically for the training phase, which continuously evaluates each layer’s parameter dynamics and learning potential to identify and prune low-contribution layers in real time. Unlike prior approaches focused on inference acceleration or backward-pass optimization, this method pioneers online forward-path compression during training, leveraging a layer-scoring mechanism to enable dynamic network scaling. Experiments on VGG and ResNet architectures across MNIST, CIFAR-10, and Imagenette datasets demonstrate over 50% reduction in training time and 17.83%–83.74% fewer forward FLOPs, all without noticeable degradation in model accuracy.