Pre-Warm: Input-Conditioned Weight Initialization for Convolutional Neural Networks

šŸ“… 2026-06-23
šŸ“ˆ Citations: 0
✨ Influential: 0
šŸ“„ PDF
šŸ¤– AI Summary
This work addresses the limitation of conventional convolutional neural network weight initialization methods, which disregard input data distribution and thereby hinder early-stage optimization. To overcome this, the authors propose Pre-Warm, a zero-training-cost, input-conditioned initialization scheme that operates prior to the first forward pass. It leverages mean-centered image patches from a single training batch, applies MiniBatchKMeans clustering, and constructs initial convolutional kernels via inverse Manhattan-space weighting. The method automatically determines nearly all hyperparameters and introduces tailored rules for predicting optimal patch counts for grayscale and color images, respectively. Evaluated across MNIST, Fashion-MNIST, CIFAR-10, SVHN, and CIFAR-100, Pre-Warm consistently outperforms Kaiming initialization (p < 0.05), achieving eight wins on SVHN and seven wins with one loss on CIFAR-100, while incurring negligible computational overhead.
šŸ“ Abstract
We introduce Pre-Warm, a simple yet effective zero-training-cost method for data-conditioned initialization of the first convolutional layer. Before the first forward pass, Pre-Warm extracts mean-centered local patches from a single training batch, clusters them with MiniBatchKMeans, applies inverse Manhattan spatial weighting, and uses the resulting centroids to initialize half of the first-layer filters (the remainder retain Kaiming initialization). We derive closed-form rules for all hyperparameters except a single insensitive scale parameter, though we derive a Kaiming parity bound on scale from patch dimensionality. For grayscale datasets we use Otsu's foreground density; for natural color images we use the mean L2 norm of mean-centered patches. Both rules accurately predict the optimal patch count observed in grid search. Across five standard benchmarks -- MNIST, Fashion-MNIST, CIFAR-10, SVHN, and CIFAR-100 -- and 8-seed paired experiments, Pre-Warm yields statistically significant accuracy improvements over standard Kaiming initialization (p < 0.05 on all datasets, p = 0.0007 on SVHN with 8/8 wins, p = 0.0033 on CIFAR-100 with 7/8 wins). The method adds negligible overhead, requires no architectural changes, and integrates into existing training pipelines with only a few lines of code. Pre-Warm demonstrates that even a lightweight, input-dependent signal can meaningfully improve optimization trajectories in modern convolutional networks.
Problem

Research questions and friction points this paper is trying to address.

weight initialization
convolutional neural networks
data-conditioned initialization
optimization trajectory
zero-training-cost
Innovation

Methods, ideas, or system contributions that make the work stand out.

data-conditioned initialization
convolutional neural networks
weight initialization
MiniBatchKMeans
zero-training-cost
šŸ”Ž Similar Papers
No similar papers found.
šŸ’¼ Related Jobs
No related jobs found.
R
Rowan Martnishn