ConsistentFeature: A Plug-and-Play Component for Neural Network Regularization

📅 2024-12-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address overfitting in over-parameterized neural networks, this paper proposes an adaptive regularization method grounded in feature consistency: it extracts features from multiple random subsets of the training set and explicitly enforces representational consistency across independent and identically distributed (i.i.d.) subsets by minimizing pairwise L2 or cosine distances among them. This is the first approach to model overfitting as *inconsistent representations on i.i.d. data subsets*, requiring no architectural or task-specific assumptions—making it plug-and-play. The method introduces zero gradient modifications and adds no extra learnable parameters. Experiments demonstrate that it significantly reduces the train–validation loss gap and improves generalization accuracy across diverse network architectures and tasks. Moreover, it exhibits strong hyperparameter robustness and negligible computational overhead.

Technology Category

Machine Learning: Learning with ManifoldsComputer Vision: Representation Learning for VisionNatural Language Processing: Learning & Optimization for NLP

Application Category

Graph Algorithms and Modeling for the Web: Graph neural networks and deep learning approaches for Web-related graphsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: User modeling for targeted and personalized online advertising
📝 Abstract
Over-parameterized neural network models often lead to significant performance discrepancies between training and test sets, a phenomenon known as overfitting. To address this, researchers have proposed numerous regularization techniques tailored to various tasks and model architectures. In this paper, we introduce a simple perspective on overfitting: models learn different representations in different i.i.d. datasets. Based on this viewpoint, we propose an adaptive method, ConsistentFeature, that regularizes the model by constraining feature differences across random subsets of the same training set. Due to minimal prior assumptions, this approach is applicable to almost any architecture and task. Our experiments show that it effectively reduces overfitting, with low sensitivity to hyperparameters and minimal computational cost. It demonstrates particularly strong memory suppression and promotes normal convergence, even when the model has already started to overfit. Even in the absence of significant overfitting, our method consistently improves accuracy and reduces validation loss.
Problem

Research questions and friction points this paper is trying to address.

Neural Networks
Overfitting
Generalization Performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

ConsistentFeature
OverfittingReduction
GeneralizationEnhancement
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
University of Chinese Academy of Science | Sichuan University
R
RuiZhe Jiang
University of Chinese Academy of Science
H
Haotian Lei
Sichuan University