Score
Design, implement, and evaluate pipelines and neural architectures that convert whole images, image regions, or frames into compact vector embeddings using convolutional and other deep feature-extraction methods, including hierarchical and geometric descriptors, and include preprocessing steps such as scaling, standardization, weighting, and fusion with other model branches. Analyze and optimize the resulting feature vectors and feature space—measuring similarity, discriminability, and robustness—and prepare embeddings for downstream uses such as retrieval, matching, and classification.
Deep convolutional neural networks (CNNs) face persistent challenges in balancing representational capacity with deployment efficiency across diverse domains—including vision, language, healthcare, and speech—especially under resource constraints and data-limited regimes. Method: This work systematically surveys CNN architectural evolution from 2015 to 2025 and proposes a novel seven-dimensional unified taxonomy (covering spatial modeling, multi-path design, dimensional expansion, attention integration, etc.), alongside a synergistic optimization framework integrating sparse convolutions, depthwise separable convolutions, and attention mechanisms. It further incorporates Fourier-based preprocessing, low-precision computation, and weight compression for lightweight deployment. Contribution/Results: We comprehensively characterize the applicability boundaries of over 100 CNN variants, quantify the trade-off between computational efficiency and representation fidelity, formalize adaptation strategies for few-shot, weakly supervised, and federated learning settings, and prospectively identify emerging directions—including CNN-Transformer hybrids, vision-language joint modeling, and generative CNNs—establishing reusable design paradigms for edge deployment and cross-domain generalization.
High barriers to adopting pre-trained models and a lack of empirical guidance for strategy selection hinder practical deployment in few-shot image classification and object detection. Method: We systematically compare linear probing versus fine-tuning across ResNet, MobileNet, and EfficientNet, and propose an end-to-end TensorFlow framework integrating multi-scale feature-space visualization (PCA, t-SNE, UMAP) to unify analysis of representation evolution. Contribution/Results: Linear probing significantly outperforms fine-tuning under extreme data scarcity (≤100 samples per class) while accelerating training by 3–5×. The framework enables high-accuracy, rapid deployment (<1 hour for fine-tuning) on standard benchmarks (ImageNet-1K, CIFAR-100), balancing beginner-friendly usability with expert-level extensibility. It bridges the gap between theoretical representation analysis and real-world engineering practice.
Traditional hand-crafted histogram features—such as Local Binary Patterns (LBP) and edge histograms—are incompatible with end-to-end deep learning due to their non-differentiability. To address this, we propose a differentiable histogram layer, enabling the first neuralization and learnability of such features. Methodologically, we design Neural Local Binary Patterns (NLBP) and Neural Edge Histogram Descriptor (NEHD) modules, integrated as differentiable statistical layers within CNNs to support gradient backpropagation and joint optimization. Our core contribution lies in unifying hand-engineered feature design with deep learning paradigms, allowing local statistical priors to be data-drivenly learned and enhanced. Extensive experiments on multiple image classification benchmarks and real-world datasets demonstrate consistent and significant performance gains, validating that neuralized histogram features substantially improve representation capability.
This work investigates the fundamental mechanisms underlying hierarchical representation learning in deep neural networks. Addressing the central question—“how do features evolve across layers”—we propose a joint quantification framework for inter-layer feature compression ratio and discriminability. We theoretically uncover, for the first time, a geometric–linear dual-rate pattern of feature evolution in deep linear networks: intra-class features contract geometrically, while inter-class discriminability increases linearly. This pattern is rigorously established under minimal norm, weight balancing, and near-low-rank assumptions, and extended to nonlinear networks via intermediate-feature modeling for multi-class classification. Numerical experiments validate its robustness across architectures and datasets. Our results provide an interpretable theoretical foundation for representation learning and yield quantitative guidance for layer selection in transfer learning and knowledge distillation.
This work addresses the challenge of detecting generative AI–produced content. We propose an unsupervised, interpretable embedding-space analysis method: semantic embeddings of text or images are extracted using pre-trained large language or multimodal models; subsequently, dimensionality reduction (e.g., PCA) uncovers an intrinsic, low-dimensional distributional shift between AI-generated and human-created samples—rendering them highly separable without supervision. This phenomenon is systematically validated for the first time and endowed with human-interpretable semantic meaning (e.g., topic coherence, syntactic redundancy). Experiments across diverse generative models—including ChatGPT, Gemini, and Stable Diffusion—demonstrate that high-accuracy separation is achieved solely from raw embeddings and unsupervised projection, without fine-tuning, labeled data, or model-specific detectors. Our approach thus significantly enhances both generalizability and interpretability of AI-content detection.
To address performance limitations in object detection and semantic segmentation under complex scenarios—including occlusion, small objects, and cross-domain generalization—this paper proposes a novel multimodal detection paradigm synergizing large language models (LLMs). Methodologically, it systematically integrates CNNs, YOLOv5/v8, and DETR architectures into an LLM-augmented inference framework, augmented by scalable data pipelines, model pruning, and quantization, and evaluated via a multi-dimensional metric system based on mAP and mIoU. Key contributions include: (1) bridging the gap between traditional feature engineering and end-to-end deep learning; (2) introducing a dynamic context enhancement mechanism tailored for challenging environments; and (3) achieving state-of-the-art accuracy-efficiency trade-offs on COCO and ADE20K. The fully open-sourced, reproducible framework significantly improves model generalizability and robustness across diverse real-world conditions.
To address slow training and excessive embedding storage overhead in ultra-fine-grained classification with millions of classes, this paper proposes a fast neural training framework based on a preconfigured latent space. The method replaces the conventional learnable classifier head with semantically orthogonal and geometrically uniform target vectors—preconstructed in a low-dimensional space using structured vector systems (e.g., Aₙ root lattices). Coupled with an encoder–ViT architecture, it enables end-to-end training without an explicit classification layer. Evaluated on ImageNet-1K and large-scale datasets containing 500K–600K classes, the approach accelerates convergence by up to 2.3× while compressing the embedding vector repository to less than 10% of that required by standard methods. It thus achieves a favorable trade-off among training efficiency, generalization performance, and deployment practicality.
Existing video compression methods are primarily optimized for human visual perception and thus often fail to preserve semantic information critical for machine vision tasks. To address this, this paper proposes a machine-vision-oriented neural preprocessing framework. It introduces a learnable preprocessor prior to standard video encoding and pioneers a differentiable virtual codec, enabling end-to-end joint optimization of preprocessing and conventional encoders (e.g., H.264/AVC) without modifying codec standards. A rate–distortion–task loss jointly optimizes bit rate, reconstruction fidelity, and downstream task performance—including object detection and action recognition. Experiments demonstrate that the framework reduces average bit rate by over 15% while maintaining or even improving task accuracy, significantly enhancing semantic fidelity and utility of compressed video for machine vision applications.
This study addresses image classification under data-scarce conditions in the context of Bangladesh by systematically comparing a lightweight custom CNN trained from scratch against widely used pretrained models—VGG-16, ResNet-50, and MobileNet—under identical experimental settings. Performance is evaluated using accuracy, precision, recall, F1-score, and computational complexity. The results demonstrate that pretrained models significantly outperform the custom architecture in terms of accuracy and convergence speed, while the custom CNN achieves competitive performance with substantially fewer parameters and lower computational overhead. These findings offer empirical evidence and practical guidance for model selection in resource-constrained scenarios where labeled data and computational resources are limited.
This study systematically compares the performance of three mainstream visual recognition strategies—custom-designed CNNs, fixed pretrained models used as feature extractors, and fine-tuned transfer learning—under real-world conditions. Through controlled experiments across five image classification datasets, the methods are comprehensively evaluated in terms of accuracy, macro F1-score, training time, and parameter count. The work provides the first empirical evidence across diverse real-world domains that transfer learning consistently achieves the best predictive performance, while custom CNNs offer a more favorable trade-off between efficiency and accuracy under computational and memory constraints. These findings offer practical guidance for model selection in applied settings where resource limitations must be balanced against performance requirements.
This work proposes a lightweight, customized CNN architecture to address the significant disparities between agricultural and urban scene images—particularly in illumination, resolution, environmental complexity, and class imbalance—and to build an efficient, robust general-purpose visual classification model. Through systematic comparisons with mainstream architectures such as ResNet-18 and VGG-16 across five heterogeneous datasets, the study evaluates the proposed model’s convergence behavior, generalization capability, and performance under both from-scratch training and transfer learning settings across varying data scales. Experimental results demonstrate that the custom CNN achieves accuracy comparable to established models while maintaining a compact footprint. Furthermore, this study provides the first systematic characterization of the practical performance boundaries of transfer learning in small-sample, highly heterogeneous scenarios, offering both theoretical insights and practical guidance for deployment under resource constraints.