Score
Designs, implements, and evaluates convolutional neural network architectures (2D and 3D) and their training pipelines to extract hierarchical, multi-scale features from spatial or spatio‑temporal data such as images, volumes, or video. Builds and optimizes convolutional layers, pooling and multi-scale modules, training regimes (minibatch optimization, data augmentation, fine‑tuning), and specialized learning procedures—including non‑negative positive‑unlabeled (NNPU) training and pseudo‑labeling—to produce and analyze models that can output pixel‑ or voxel‑wise, explainable quality or localization maps for detecting localized degradation or other fine‑grained signals.
This study investigates the selection of efficient convolutional neural network (CNN) architectures for image classification and object detection under varying task complexity and resource constraints. Through systematic comparisons across five real-world datasets—including binary classification, fine-grained multiclass classification, and object detection tasks—the work evaluates custom CNNs, deep residual networks, and transfer learning models, analyzing the impact of key architectural factors such as network depth and residual connections. The results demonstrate that deeper architectures significantly improve accuracy in fine-grained classification, whereas lightweight pretrained models offer superior efficiency for simpler binary classification tasks. Furthermore, the proposed custom CNN is successfully extended to detect illegally operating tricycles in traffic scenarios, confirming its practical effectiveness in real-world applications.
This work proposes a lightweight, customized CNN architecture to address the significant disparities between agricultural and urban scene images—particularly in illumination, resolution, environmental complexity, and class imbalance—and to build an efficient, robust general-purpose visual classification model. Through systematic comparisons with mainstream architectures such as ResNet-18 and VGG-16 across five heterogeneous datasets, the study evaluates the proposed model’s convergence behavior, generalization capability, and performance under both from-scratch training and transfer learning settings across varying data scales. Experimental results demonstrate that the custom CNN achieves accuracy comparable to established models while maintaining a compact footprint. Furthermore, this study provides the first systematic characterization of the practical performance boundaries of transfer learning in small-sample, highly heterogeneous scenarios, offering both theoretical insights and practical guidance for deployment under resource constraints.
High barriers to adopting pre-trained models and a lack of empirical guidance for strategy selection hinder practical deployment in few-shot image classification and object detection. Method: We systematically compare linear probing versus fine-tuning across ResNet, MobileNet, and EfficientNet, and propose an end-to-end TensorFlow framework integrating multi-scale feature-space visualization (PCA, t-SNE, UMAP) to unify analysis of representation evolution. Contribution/Results: Linear probing significantly outperforms fine-tuning under extreme data scarcity (≤100 samples per class) while accelerating training by 3–5×. The framework enables high-accuracy, rapid deployment (<1 hour for fine-tuning) on standard benchmarks (ImageNet-1K, CIFAR-100), balancing beginner-friendly usability with expert-level extensibility. It bridges the gap between theoretical representation analysis and real-world engineering practice.
Conventional CNNs struggle to model high-order pixel-wise correlations in images. Method: Inspired by nonlinear processing mechanisms in biological vision, we propose a learnable high-order Volterra convolution module that explicitly models multiplicative interactions among pixels. This is the first incorporation of biologically plausible high-order nonlinear convolution into deep visual models, supporting dynamic learning of the optimal expansion order—empirically found to be 3–4, aligning with statistical properties of natural images. Contribution/Results: Through representational similarity analysis (RSA), systematic perturbation studies, and evaluation across multiple datasets (MNIST–Imagenette), our method achieves significant performance gains over standard CNNs on CIFAR-10/100. It reveals order-specific encoding of distinct visual information subdimensions and characterizes hierarchical differences in representational geometry across network layers—establishing a novel paradigm for interpretable, biologically grounded visual modeling.
Traditional hand-crafted histogram features—such as Local Binary Patterns (LBP) and edge histograms—are incompatible with end-to-end deep learning due to their non-differentiability. To address this, we propose a differentiable histogram layer, enabling the first neuralization and learnability of such features. Methodologically, we design Neural Local Binary Patterns (NLBP) and Neural Edge Histogram Descriptor (NEHD) modules, integrated as differentiable statistical layers within CNNs to support gradient backpropagation and joint optimization. Our core contribution lies in unifying hand-engineered feature design with deep learning paradigms, allowing local statistical priors to be data-drivenly learned and enhanced. Extensive experiments on multiple image classification benchmarks and real-world datasets demonstrate consistent and significant performance gains, validating that neuralized histogram features substantially improve representation capability.
This study investigates the impact of convolutional neural network (CNN) architecture design on image classification performance across agricultural and urban domains. To this end, the authors propose a customized CNN—CustomCNN—that integrates residual connections, Squeeze-and-Excitation attention mechanisms, and a progressive channel scaling strategy, along with Kaiming initialization to enhance representational capacity and training efficiency. Experimental results on five publicly available datasets demonstrate that the proposed model achieves classification performance comparable to that of state-of-the-art CNNs while maintaining computational efficiency. These findings underscore the effectiveness and potential of domain-informed architectural design for visual applications in smart cities and precision agriculture.
This study systematically compares the performance of three mainstream visual recognition strategies—custom-designed CNNs, fixed pretrained models used as feature extractors, and fine-tuned transfer learning—under real-world conditions. Through controlled experiments across five image classification datasets, the methods are comprehensively evaluated in terms of accuracy, macro F1-score, training time, and parameter count. The work provides the first empirical evidence across diverse real-world domains that transfer learning consistently achieves the best predictive performance, while custom CNNs offer a more favorable trade-off between efficiency and accuracy under computational and memory constraints. These findings offer practical guidance for model selection in applied settings where resource limitations must be balanced against performance requirements.
Standard convolutions, due to their fixed structure, linearity, and reliance on local averaging, struggle to capture complex image characteristics such as low-rank structures, adaptive basis representations, and non-uniform spatial dependencies. This work proposes a unified taxonomy encompassing five classes of structured operators—decomposition-based, adaptive weighting, basis-adaptive, integral/kernel-based, and attention-based—and systematically analyzes their differences along key dimensions including locality, linearity, and equivariance. By leveraging techniques such as singular value/tensor decomposition, content-adaptive weighting, learnable analysis bases, position-dependent nonlinear kernels, and attention mechanisms, the study comprehensively evaluates the performance of these operators across image-to-image and image-to-label tasks. The findings clarify the respective strengths and limitations of each operator class, offering both theoretical insights and practical guidance for future research.
This study investigates the generalization capability and overfitting behavior of neural networks on the CIFAR-10 image classification task. By constructing and comparing a fully connected network with a convolutional architecture comprising six convolutional layers and three max-pooling layers, the work implements a complete pipeline encompassing data preprocessing (normalization and one-hot encoding), training (using the Adam optimizer with mini-batches), and validation. After ten training epochs, the model achieves a validation accuracy of 74.77% and clearly exhibits the hallmark overfitting pattern: training loss continues to decrease while validation loss begins to rise. The findings underscore the distinction between representation learning and mere memorization, offering a reproducible benchmark framework that can inform the development of regularization techniques, data augmentation strategies, and educational experimentation.
This study addresses image classification under data-scarce conditions in the context of Bangladesh by systematically comparing a lightweight custom CNN trained from scratch against widely used pretrained models—VGG-16, ResNet-50, and MobileNet—under identical experimental settings. Performance is evaluated using accuracy, precision, recall, F1-score, and computational complexity. The results demonstrate that pretrained models significantly outperform the custom architecture in terms of accuracy and convergence speed, while the custom CNN achieves competitive performance with substantially fewer parameters and lower computational overhead. These findings offer empirical evidence and practical guidance for model selection in resource-constrained scenarios where labeled data and computational resources are limited.