Impact of Noisy Supervision in Foundation Model Learning

📅 2024-03-11
📈 Citations: 3
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work investigates how label noise in pretraining data affects the generalization of foundation models, revealing that while such noise may improve in-distribution (ID) performance, it inevitably degrades out-of-distribution (OOD) generalization—primarily by distorting the feature space geometry. To address this, the authors introduce “noisy-model tuning” as a novel paradigm and propose NMTune, a generic, parameter-efficient feature-space calibration method applicable to both white-box and black-box models. Extensive evaluation across synthetic and real-world noisy datasets—including ImageNet-1K, YFCC15M, and CC12M—covers diverse pretraining paradigms (fully supervised and vision-language contrastive), model architectures, downstream tasks, and tuning strategies. Results demonstrate that NMTune consistently mitigates noise-induced degradation, significantly improving OOD generalization across vision and language models—including proprietary API-based models—without dependence on model scale or task specificity.

Technology Category

Machine Learning: Calibration & Uncertainty QuantificationComputer Vision: Diffusion Models for VisionNatural Language Processing: (Large) Language Models

Application Category

Web Mining and Content Analysis: Large pretrained models with web dataGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsUser Modeling, Personalization and Recommendation: Attacks and countermeasures in recommendation systems
📝 Abstract
Foundation models are usually pre-trained on large-scale datasets and then adapted to downstream tasks through tuning. However, the large-scale pre-training datasets, often inaccessible or too expensive to handle, can contain label noise that may adversely affect the generalization of the model and pose unexpected risks. This paper stands out as the first work to comprehensively understand and analyze the nature of noise in pre-training datasets and then effectively mitigate its impacts on downstream tasks. Specifically, through extensive experiments of fully-supervised and image-text contrastive pre-training on synthetic noisy ImageNet-1K, YFCC15M, and CC12M datasets, we demonstrate that, while slight noise in pre-training can benefit in-domain (ID) performance, where the training and testing data share a similar distribution, it always deteriorates out-of-domain (OOD) performance, where training and testing distributions are significantly different. These observations are agnostic to scales of pre-training datasets, pre-training noise types, model architectures, pre-training objectives, downstream tuning methods, and downstream applications. We empirically ascertain that the reason behind this is that the pre-training noise shapes the feature space differently. We then propose a tuning method (NMTune) to affine the feature space to mitigate the malignant effect of noise and improve generalization, which is applicable in both parameter-efficient and black-box tuning manners. We additionally conduct extensive experiments on popular vision and language models, including APIs, which are supervised and self-supervised pre-trained on realistic noisy data for evaluation. Our analysis and results demonstrate the importance of this novel and fundamental research direction, which we term as Noisy Model Learning.
Problem

Research questions and friction points this paper is trying to address.

Analyzes impact of label noise in pre-training datasets on model generalization.
Proposes NMTune method to mitigate noise effects on downstream tasks.
Explores noise influence across various datasets, models, and tuning methods.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Analyzes noise impact in pre-training datasets
Proposes NMTune for noise mitigation
Improves generalization across diverse applications
🔎 Similar Papers
No similar papers found.