contrastive denoising training

Design and implement training procedures that corrupt targets or inputs and train models to reconstruct the clean targets while simultaneously applying contrastive losses to distinguish true targets from near-duplicates or corrupted alternatives. Build loss functions, sampling/corruption schedules, and evaluation analyses that penalize duplicate or spurious predictions, stabilize multi-step denoising processes, and improve convergence and robustness to label noise and matching ambiguity.

contrastivedenoisingtraining

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.35
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Corruptions of Supervised Learning Problems: Typology and Mitigations

Jul 17, 2023
LI
Laura Iacovissi
🏛️ University of Tübingen | Tübingen AI Center

Existing data corruption studies are fragmented across specific scenarios, lacking a unified theoretical framework and systematic mitigation strategies. Method: We propose the first general corruption modeling framework based on Markov kernels, formalizing corruption as arbitrary modifications to the data distribution, hypothesis class, or loss function. We establish a provably complete taxonomy—distinguishing, for the first time, label corruption (affecting only the loss) from attribute or joint corruption (simultaneously affecting both the hypothesis class and the loss). Building on this, we introduce a generalized loss correction paradigm, deriving provably effective correction formulas for attribute and joint corruption under weaker assumptions than conventional approaches. Contribution/Results: Our framework unifies disparate corruption models and terminologies, providing a rigorous foundation for robustness analysis and algorithm design in supervised learning. It enables principled treatment of previously isolated corruption types and advances theoretical understanding of learning under distributional and structural perturbations.

Analyzes corruption impacts on loss functions and hypothesis classesDevelops a unified theory of corruption in supervised learning problemsProposes generalized mitigation methods for diverse corruption types

This study investigates the joint impact of data missingness and noise on machine learning performance. We systematically quantify trade-offs among data quality, volume, and imputation strategies across two representative scenarios: NLP supervised learning (BERT) and traffic signal control via reinforcement learning (PPO). Methodologically, we propose a novel “Performance Degradation Index Model under Data Corruption” and identify that only 30% of critical data governs overall model performance. We introduce the concepts of “Imputation Advantage Angle” and “Imputation Disadvantage Edge,” and— for the first time—categorize learning tasks into noise-sensitive versus noise-insensitive classes. Results show that noise degrades performance more severely than missingness; imputation efficacy critically depends on alignment between imputation accuracy and data corruption rate; and merely scaling data volume only mitigates—not eliminates—corruption effects, with diminishing marginal returns intensifying as corruption worsens.

Data Imputation StrategiesData QualityMachine Learning

This study investigates the robust learning mechanisms of neural networks when input data are severely corrupted by attribute noise—such as additive or replacement noise—while labels remain intact. Through experiments with multilayer perceptrons, mean-field analysis of infinite-width networks, and prototype-based classification theory, the work reveals for the first time that networks consistently adopt a nearest-class-mean decision rule even when over 90% of the input features are corrupted. This behavior is shown to be universal across network depth, activation functions, and noise distributions. Empirical results demonstrate that finite-width networks closely align with theoretical predictions, significantly outperforming random guessing under extreme noise conditions. These findings establish an interpretable and analytically tractable foundation for understanding robustness in deep learning.

attribute noiseclassificationinput corruption

This work addresses a critical yet overlooked issue in adaptive data cleaning: fluctuating sample removal counts caused by varying partition granularities introduce budget confounding bias, leading to spurious performance gains falsely attributed to contamination identification capability. To enable fair evaluation, the authors propose an operating-point-matched assessment framework that aligns removal budgets with recall rates and incorporates threshold-agnostic metrics (AUROC and AUPRC). They systematically uncover and resolve this budget confounding problem for the first time, introducing a multi-cue adaptive cleaner—integrating learning difficulty reweighting, Euclidean distance guidance, and fine-grained partitioning—and a false positive decomposition analysis. Experiments on CIFAR-10 and ImageNet-100 reveal that most existing methods lose their apparent advantage under matched operating points, demonstrating genuine efficacy only under low contamination rates or high-recall, heavily corrupted scenarios.

adaptive data cleaningcorruption discriminationevaluation bias

Alignment Calibration: Machine Unlearning for Contrastive Learning under Auditing

Jun 05, 2024
YW
Yihan Wang
🏛️ Academy of Mathematics and Systems Science | Chinese Academy of Sciences | University of Chinese Academy of Sciences | University of Waterloo | Vector Institute | Huawei Noah’s Ark Lab | CISPA

This work presents the first systematic study of machine unlearning for contrastive learning (CL) models, identifying two critical gaps: the failure of existing unlearning methods under the CL paradigm and the absence of appropriate evaluation protocols. To address these, we propose MUC, a CL-specific unlearning framework whose core innovation is Alignment Calibration—a principled, auditable optimization objective leveraging CL’s alignment property. MUC achieves controllable, representation-level forgetting after sensitive data removal via contrastive loss recalibration and embedding-space alignment constraints. The framework is architecture-agnostic, supporting mainstream CL models including SimCLR, MoCo, and CLIP, and introduces novel audit metrics enabling black-box verification and visual interpretability. Extensive experiments demonstrate that MUC consistently approaches full retraining performance across multiple benchmarks, substantially outperforming prior unlearning methods. Notably, MUC establishes the first verifiable and interpretable data unlearning capability for CL models.

Addressing limitations in current unlearning validation approachesEvaluating unlearning effects with black-box methodsMachine unlearning for contrastive learning models

Latest Papers

What's happening recently
View more

This work addresses a critical limitation in existing reconstruction-based methods for learning with noisy labels: their tendency to jointly assess the reliability of observed labels and pseudo-targets, which often leads to unreliable signals substituting one another and hinders effective denoising. To overcome this, the paper proposes TRACE, a novel framework that decouples the reliability evaluation of these two sources for the first time. Specifically, it evaluates observed labels through loss fitting, shallow-feature relational stability, and prediction consistency, while assessing pseudo-targets via model confidence. These independent reliability estimates are then used to separately govern label correction and sample reweighting. By preventing error propagation, TRACE generates more trustworthy pseudo-supervision, significantly outperforming current reconstruction-based approaches across multiple synthetic and real-world noisy benchmarks, and thereby enhancing model robustness and generalization.

label-noise learningnoisy labelspseudo targets

This work proposes a general-purpose loss function based on a Bayesian latent-variable switching mixture model that simultaneously enables robust learning and unsupervised identification of corrupted labels during training. While existing robust loss functions mitigate the adverse effects of label noise, they lack the ability to explicitly detect contaminated samples. The proposed approach uniquely unifies robust loss with interpretable posterior probabilities of label corruption, incorporates input-dependent priors to capture the spatial locality of noise, and naturally induces an Occam’s razor regularization via marginal likelihood to prevent over-flagging. Experiments on CIFAR-10 under asymmetric label noise (corruption rates 0.2–0.6) demonstrate that the method accurately recovers the underlying corruption structure, effectively distinguishes clean from noisy samples, identifies the direction of label flips, and significantly outperforms four state-of-the-art robust loss baselines.

anomaly detectionBayesian modelinglabel contamination

This work systematically investigates the impact of loss weighting strategies and output parameterizations on model performance in flow matching. Through numerical experiments on both synthetic data with controllable geometric structures and real-world images, the study disentangles their interaction effects across varying data manifold dimensions, model architectures, and dataset scales, using PSNR and FID as evaluation metrics. The analysis reveals, for the first time, how the optimal choice of loss weighting and parameterization depends critically on the intrinsic structure of the data. Building on these insights, the authors formulate practical design principles that substantially improve denoising accuracy and generation quality.

denoisingflow matchinggenerative models

Hot Scholars

WB

Wolfram Burgard

Professor of Computer Science, University of Technology Nuremberg
RoboticsArtificial IntelligenceAIMachine Learning
TK

Tassilo Klein

Principal Scientist @ SAP AI Research
NLPTable Representation LearningMedical ImagingComputer Vision
CS

Chandan Singh

Senior researcher, Microsoft research
🔍 Interpretability🤖 Foundation models🧠 Neuroscience🌳 Transparent models
JB

James Bailey

Professor, School of Computing and Information Systems, University of Melbourne
machine learningartificial intelligencedata mining
JL

Jundong Li

Associate Professor, University of Virginia
AIMachine LearningData MiningGraph Learning