Score
Design and implement models that adapt a pretrained convolutional neural network by replacing or modifying the network head and updating selected weights (e.g., retraining output layers and optionally some convolutional blocks) on a new labeled dataset to create a classifier or regressor for a target task. This work includes choosing which layers to freeze, setting per-layer learning rates and regularization, and handling limited-data regimes to transfer learned features from the source model to the new task.
Deep models for image classification heavily rely on large-scale labeled data, yet real-world scenarios often suffer from data scarcity, making transfer learning a critical solution—yet systematic surveys and theoretical frameworks remain lacking. This paper proposes the first unified taxonomy for deep transfer learning in image classification, formally defining the problem, identifying core challenges (e.g., source/target domain shift, limited target-sample size), and characterizing failure boundaries. It systematically integrates both CNN- and Transformer-based architectures, covering major paradigms: feature extraction, fine-tuning, domain adaptation, and meta-transfer learning. Through cross-paradigm comparative analysis, it uncovers intrinsic relationships between model performance, data characteristics (e.g., domain gap, sample scale), and architectural/algorithmic choices. The study clarifies current research gaps, establishes necessary conditions for effective transfer, and delivers a reusable methodological framework for few-shot and cross-domain image classification.
This work addresses the degradation of neural plasticity in pretrained models during transfer to downstream tasks, a phenomenon often caused by weight saturation that impairs adaptation to atypical data. To mitigate this issue, the study introduces— for the first time—a systematic neuroplasticity restoration mechanism in transfer learning through a lightweight, architecture-agnostic targeted weight reinitialization strategy applied prior to fine-tuning. The proposed method effectively alleviates weight saturation without altering the standard training pipeline and is compatible with both convolutional neural networks (CNNs) and Vision Transformers (ViTs). Extensive experiments demonstrate that this approach consistently accelerates convergence and improves final accuracy across multiple image classification benchmarks, all while incurring negligible computational overhead.
Deploying CNNs on resource-constrained devices faces challenges of high computational overhead and inflexible architectures. Method: This paper proposes an elastic CNN architecture enabling zero-shot, runtime adaptation to multiple granularities of computational complexity without fine-tuning. It introduces a novel pruning-growth co-design paradigm for nested subnetwork construction, integrating structured pruning with dynamic subnet reconfiguration to enable seamless switching between compact and full configurations within a single model. Contribution/Results: Evaluated on VGG-16, AlexNet, and ResNet across CIFAR-10 and Imagenette, the approach reduces computation by 40% with <1.2% accuracy degradation—and in some cases surpasses baseline accuracy. To our knowledge, this is the first work achieving real-time, training-free capacity adaptation, significantly enhancing deployment flexibility and energy efficiency of edge AI models.
Traditionally “untrainable” architectures—such as fully connected networks, residual-free CNNs, and vanilla RNNs—suffer from severe overfitting or underfitting due to insufficient inductive bias. To address this, we propose a representation alignment guidance mechanism that transfers architectural priors from high-performance “guide” networks (e.g., ResNet, Transformer) to target networks via differentiable neural distance functions, enforcing layer-wise representation alignment. Crucially, the guide network remains frozen while the target network is jointly optimized for both task loss and alignment loss. This work introduces the first quantifiable, differentiable framework for architectural prior transfer, providing a mathematically grounded, optimization-compatible tool for neural architecture design. Experiments demonstrate substantial improvements: fully connected networks achieve markedly enhanced visual generalization; plain CNNs approach ResNet-level performance; the performance gap between vanilla RNNs and Transformers narrows significantly; and remarkably, Transformers exhibit improved accuracy on RNN-favored tasks—a reverse enhancement effect.
This work addresses the limitation of conventional convolutional neural network weight initialization methods, which disregard input data distribution and thereby hinder early-stage optimization. To overcome this, the authors propose Pre-Warm, a zero-training-cost, input-conditioned initialization scheme that operates prior to the first forward pass. It leverages mean-centered image patches from a single training batch, applies MiniBatchKMeans clustering, and constructs initial convolutional kernels via inverse Manhattan-space weighting. The method automatically determines nearly all hyperparameters and introduces tailored rules for predicting optimal patch counts for grayscale and color images, respectively. Evaluated across MNIST, Fashion-MNIST, CIFAR-10, SVHN, and CIFAR-100, Pre-Warm consistently outperforms Kaiming initialization (p < 0.05), achieving eight wins on SVHN and seven wins with one loss on CIFAR-100, while incurring negligible computational overhead.
High barriers to adopting pre-trained models and a lack of empirical guidance for strategy selection hinder practical deployment in few-shot image classification and object detection. Method: We systematically compare linear probing versus fine-tuning across ResNet, MobileNet, and EfficientNet, and propose an end-to-end TensorFlow framework integrating multi-scale feature-space visualization (PCA, t-SNE, UMAP) to unify analysis of representation evolution. Contribution/Results: Linear probing significantly outperforms fine-tuning under extreme data scarcity (≤100 samples per class) while accelerating training by 3–5×. The framework enables high-accuracy, rapid deployment (<1 hour for fine-tuning) on standard benchmarks (ImageNet-1K, CIFAR-100), balancing beginner-friendly usability with expert-level extensibility. It bridges the gap between theoretical representation analysis and real-world engineering practice.
This work addresses catastrophic forgetting in fine-tuning pretrained models, where newly acquired knowledge overwrites previously learned information. To mitigate this issue, the authors propose a function-preserving model expansion approach that mathematically duplicates and scales parameters of selected Transformer submodules during initialization. This technique enables stable training and faithful retention of original model capabilities without altering the initial functionality. By circumventing the traditional trade-off between plasticity and stability, the method achieves performance comparable to full fine-tuning while expanding only a minimal number of layers. Consequently, it fully preserves the model’s original knowledge and substantially reduces computational overhead.
Existing theoretical frameworks struggle to explain why larger-scale pre-trained models substantially reduce sample complexity on downstream tasks. This work proposes a novel theoretical framework—termed “caulking”—inspired by parameter-efficient fine-tuning methods such as adapters, low-rank adaptation, and partial fine-tuning. It establishes, for the first time, a provable relationship between the scale of pre-trained models and the sample complexity of downstream tasks. By rigorously linking stronger pre-training capabilities to reduced data requirements in transfer learning, this study not only addresses a critical gap in current theoretical understanding but also provides a solid foundation for empirically observed scaling laws, demonstrating that enhanced pre-training capacity can significantly decrease the number of samples needed for effective downstream adaptation.
This work addresses the high computational cost of forward propagation during the training of deep convolutional neural networks, which significantly limits training efficiency. The authors propose a dynamic layer pruning method tailored specifically for the training phase, which continuously evaluates each layer’s parameter dynamics and learning potential to identify and prune low-contribution layers in real time. Unlike prior approaches focused on inference acceleration or backward-pass optimization, this method pioneers online forward-path compression during training, leveraging a layer-scoring mechanism to enable dynamic network scaling. Experiments on VGG and ResNet architectures across MNIST, CIFAR-10, and Imagenette datasets demonstrate over 50% reduction in training time and 17.83%–83.74% fewer forward FLOPs, all without noticeable degradation in model accuracy.
This study addresses the challenge of selecting suitable ImageNet-pretrained models for image classification tasks in target domains by introducing a multidimensional evaluation framework. The authors systematically fine-tune the output layers and general parameters of eleven pretrained models across five diverse datasets, evaluating their performance under both single-run and multiple-run training settings. Through comprehensive assessment of accuracy, accuracy density, training time, and model size, the work quantifies the cross-domain transferability differences among pretrained models, revealing consistent patterns in how model characteristics align with task-specific requirements. These findings provide empirical evidence and practical guidelines for informed model selection in real-world applications.
This work addresses the inflexibility of conventional pre-trained models, whose fixed sizes hinder adaptation to downstream tasks requiring varying model scales. The authors propose a novel pre-training paradigm based on structured constraints, introducing Kronecker factorization into pre-training for the first time. This approach decouples model weights into a scale-invariant, reusable weight template and a lightweight, data-driven weight scaler, framing variable-scale model initialization as a multi-task adaptation problem. The method supports arbitrary depths and widths in both Transformer and CNN architectures, significantly accelerating convergence and improving performance across diverse tasks—including image classification, generation, and embodied control—thereby enabling efficient and flexible model deployment.