Score
Designs and implements image classification systems that extract deep convolutional neural network embeddings (e.g., CNN features or EfficientNet embeddings) and train classical ensemble classifiers such as random forests on those embeddings. Builds and evaluates the end-to-end pipeline — feature extraction, classifier training and tuning — and analyzes accuracy, inference latency, and trade-offs between embedding choice and ensemble configuration.
Deep models for image classification heavily rely on large-scale labeled data, yet real-world scenarios often suffer from data scarcity, making transfer learning a critical solution—yet systematic surveys and theoretical frameworks remain lacking. This paper proposes the first unified taxonomy for deep transfer learning in image classification, formally defining the problem, identifying core challenges (e.g., source/target domain shift, limited target-sample size), and characterizing failure boundaries. It systematically integrates both CNN- and Transformer-based architectures, covering major paradigms: feature extraction, fine-tuning, domain adaptation, and meta-transfer learning. Through cross-paradigm comparative analysis, it uncovers intrinsic relationships between model performance, data characteristics (e.g., domain gap, sample scale), and architectural/algorithmic choices. The study clarifies current research gaps, establishes necessary conditions for effective transfer, and delivers a reusable methodological framework for few-shot and cross-domain image classification.
Single-architectural vision models face inherent performance bottlenecks in image classification. Method: This paper proposes a cross-paradigm ensemble framework that preserves architectural integrity by integrating three heterogeneous architectures—CNN (ResNet), MLP-Mixer, and Vision Transformer—via weighted averaging and majority voting, without architectural modification or feature-space alignment. Crucially, the framework implicitly isolates their respective feature spaces while enabling synergistic gains. Contribution/Results: It introduces the first complementary analysis paradigm grounded in architectural orthogonality. Evaluated end-to-end on ImageNet, the ensemble surpasses prior single-model SOTA in top-1 accuracy while reducing overall inference latency—establishing a new benchmark for efficient, high-accuracy image classification.
High barriers to adopting pre-trained models and a lack of empirical guidance for strategy selection hinder practical deployment in few-shot image classification and object detection. Method: We systematically compare linear probing versus fine-tuning across ResNet, MobileNet, and EfficientNet, and propose an end-to-end TensorFlow framework integrating multi-scale feature-space visualization (PCA, t-SNE, UMAP) to unify analysis of representation evolution. Contribution/Results: Linear probing significantly outperforms fine-tuning under extreme data scarcity (≤100 samples per class) while accelerating training by 3–5×. The framework enables high-accuracy, rapid deployment (<1 hour for fine-tuning) on standard benchmarks (ImageNet-1K, CIFAR-100), balancing beginner-friendly usability with expert-level extensibility. It bridges the gap between theoretical representation analysis and real-world engineering practice.
Accurate shape-aware classification of single-object images—particularly in e-commerce settings—remains challenging due to the semantic gap between low-level geometric representations and high-level semantics. Method: We propose a hierarchical shape-feature classification framework that bridges this gap by integrating image segmentation with object recognition pre- and post-processing. Leveraging shape features, we construct four distinct classifiers based on Bayesian networks, random forests, Bagging, and voting ensembles, and conduct the first systematic evaluation of ensemble strategies for single-object classification. Experiments employ 10-fold cross-validation on Amazon and Google single-object datasets. Results/Contributions: (1) Bagging achieves 99% classification accuracy—significantly outperforming baselines—validating the effectiveness of synergistic shape-feature and Bagging modeling; (2) ensemble learning is empirically established as superior for fine-grained single-object classification; (3) the framework supports scalable automated annotation and cross-platform image retrieval.
This work challenges the prevailing assumption that dataset bias has been mitigated in modern vision models, systematically investigating their ability to discriminate image provenance. Using three large-scale open datasets—YFCC, Conceptual Captions (CC), and DataComp—we conduct three-way classification experiments with state-of-the-art pre-trained vision models and complement them with feature transferability and generalization analyses. Results show that models achieve 84.7% accuracy in identifying the source dataset on held-out validation sets, demonstrating persistent and substantial dataset-level bias. Crucially, the discriminative features exhibit semantic coherence and cross-task transferability, providing the first empirical evidence that large models learn generalizable semantic patterns—not mere memorization. This reveals latent systemic biases in current dataset curation and model evaluation practices, offering new perspectives for bias modeling, dataset auditing, and robust generalization research.
Traditional hand-crafted histogram features—such as Local Binary Patterns (LBP) and edge histograms—are incompatible with end-to-end deep learning due to their non-differentiability. To address this, we propose a differentiable histogram layer, enabling the first neuralization and learnability of such features. Methodologically, we design Neural Local Binary Patterns (NLBP) and Neural Edge Histogram Descriptor (NEHD) modules, integrated as differentiable statistical layers within CNNs to support gradient backpropagation and joint optimization. Our core contribution lies in unifying hand-engineered feature design with deep learning paradigms, allowing local statistical priors to be data-drivenly learned and enhanced. Extensive experiments on multiple image classification benchmarks and real-world datasets demonstrate consistent and significant performance gains, validating that neuralized histogram features substantially improve representation capability.
This work proposes a performance optimization approach for CIFAR-10 image classification that avoids indiscriminately increasing model complexity. Through systematic ablation studies, the authors evaluate 17 training and architectural enhancements—including learning rate scheduling, Dropout, pooling strategies, and configurations of network depth and filter counts—to identify the most effective components. These are then integrated into a weighted ensemble model to improve generalization. Emphasizing empirically driven fine-tuning over mere scaling of model size, the method achieves 89.23% accuracy on the full CIFAR-10 test set, demonstrating the efficacy of strategic component selection and ensemble learning within lightweight convolutional neural networks.
This study addresses image classification under data-scarce conditions in the context of Bangladesh by systematically comparing a lightweight custom CNN trained from scratch against widely used pretrained models—VGG-16, ResNet-50, and MobileNet—under identical experimental settings. Performance is evaluated using accuracy, precision, recall, F1-score, and computational complexity. The results demonstrate that pretrained models significantly outperform the custom architecture in terms of accuracy and convergence speed, while the custom CNN achieves competitive performance with substantially fewer parameters and lower computational overhead. These findings offer empirical evidence and practical guidance for model selection in resource-constrained scenarios where labeled data and computational resources are limited.
This study investigates the generalization capability and overfitting behavior of neural networks on the CIFAR-10 image classification task. By constructing and comparing a fully connected network with a convolutional architecture comprising six convolutional layers and three max-pooling layers, the work implements a complete pipeline encompassing data preprocessing (normalization and one-hot encoding), training (using the Adam optimizer with mini-batches), and validation. After ten training epochs, the model achieves a validation accuracy of 74.77% and clearly exhibits the hallmark overfitting pattern: training loss continues to decrease while validation loss begins to rise. The findings underscore the distinction between representation learning and mere memorization, offering a reproducible benchmark framework that can inform the development of regularization techniques, data augmentation strategies, and educational experimentation.
This study addresses the challenge of fine-grained classification within the ambiguous “Other” phenotype category in zebrafish embryo images by proposing a two-stage hierarchical ensemble approach. In the first stage, a four-class model identifies major phenotypes; in the second stage, three specialized ensemble architectures further refine the “Other” class, with Setup 2—incorporating a multi-label classifier—demonstrating superior performance and better class balance. Experiments conducted using ResNet18, Vision Transformer (ViT), and ConvNeXt backbones reveal that ConvNeXt substantially enhances feature representation and consistently achieves the best results across all configurations, thereby validating the efficacy and advancement of the proposed hierarchical ensemble strategy.
This work proposes a lightweight, customized CNN architecture to address the significant disparities between agricultural and urban scene images—particularly in illumination, resolution, environmental complexity, and class imbalance—and to build an efficient, robust general-purpose visual classification model. Through systematic comparisons with mainstream architectures such as ResNet-18 and VGG-16 across five heterogeneous datasets, the study evaluates the proposed model’s convergence behavior, generalization capability, and performance under both from-scratch training and transfer learning settings across varying data scales. Experimental results demonstrate that the custom CNN achieves accuracy comparable to established models while maintaining a compact footprint. Furthermore, this study provides the first systematic characterization of the practical performance boundaries of transfer learning in small-sample, highly heterogeneous scenarios, offering both theoretical insights and practical guidance for deployment under resource constraints.