Score
Design and implement training procedures that corrupt targets or inputs and train models to reconstruct the clean targets while simultaneously applying contrastive losses to distinguish true targets from near-duplicates or corrupted alternatives. Build loss functions, sampling/corruption schedules, and evaluation analyses that penalize duplicate or spurious predictions, stabilize multi-step denoising processes, and improve convergence and robustness to label noise and matching ambiguity.
Existing data corruption studies are fragmented across specific scenarios, lacking a unified theoretical framework and systematic mitigation strategies. Method: We propose the first general corruption modeling framework based on Markov kernels, formalizing corruption as arbitrary modifications to the data distribution, hypothesis class, or loss function. We establish a provably complete taxonomy—distinguishing, for the first time, label corruption (affecting only the loss) from attribute or joint corruption (simultaneously affecting both the hypothesis class and the loss). Building on this, we introduce a generalized loss correction paradigm, deriving provably effective correction formulas for attribute and joint corruption under weaker assumptions than conventional approaches. Contribution/Results: Our framework unifies disparate corruption models and terminologies, providing a rigorous foundation for robustness analysis and algorithm design in supervised learning. It enables principled treatment of previously isolated corruption types and advances theoretical understanding of learning under distributional and structural perturbations.
This study investigates the joint impact of data missingness and noise on machine learning performance. We systematically quantify trade-offs among data quality, volume, and imputation strategies across two representative scenarios: NLP supervised learning (BERT) and traffic signal control via reinforcement learning (PPO). Methodologically, we propose a novel “Performance Degradation Index Model under Data Corruption” and identify that only 30% of critical data governs overall model performance. We introduce the concepts of “Imputation Advantage Angle” and “Imputation Disadvantage Edge,” and— for the first time—categorize learning tasks into noise-sensitive versus noise-insensitive classes. Results show that noise degrades performance more severely than missingness; imputation efficacy critically depends on alignment between imputation accuracy and data corruption rate; and merely scaling data volume only mitigates—not eliminates—corruption effects, with diminishing marginal returns intensifying as corruption worsens.
This study investigates the robust learning mechanisms of neural networks when input data are severely corrupted by attribute noise—such as additive or replacement noise—while labels remain intact. Through experiments with multilayer perceptrons, mean-field analysis of infinite-width networks, and prototype-based classification theory, the work reveals for the first time that networks consistently adopt a nearest-class-mean decision rule even when over 90% of the input features are corrupted. This behavior is shown to be universal across network depth, activation functions, and noise distributions. Empirical results demonstrate that finite-width networks closely align with theoretical predictions, significantly outperforming random guessing under extreme noise conditions. These findings establish an interpretable and analytically tractable foundation for understanding robustness in deep learning.
This work addresses a critical yet overlooked issue in adaptive data cleaning: fluctuating sample removal counts caused by varying partition granularities introduce budget confounding bias, leading to spurious performance gains falsely attributed to contamination identification capability. To enable fair evaluation, the authors propose an operating-point-matched assessment framework that aligns removal budgets with recall rates and incorporates threshold-agnostic metrics (AUROC and AUPRC). They systematically uncover and resolve this budget confounding problem for the first time, introducing a multi-cue adaptive cleaner—integrating learning difficulty reweighting, Euclidean distance guidance, and fine-grained partitioning—and a false positive decomposition analysis. Experiments on CIFAR-10 and ImageNet-100 reveal that most existing methods lose their apparent advantage under matched operating points, demonstrating genuine efficacy only under low contamination rates or high-recall, heavily corrupted scenarios.
This work presents the first systematic study of machine unlearning for contrastive learning (CL) models, identifying two critical gaps: the failure of existing unlearning methods under the CL paradigm and the absence of appropriate evaluation protocols. To address these, we propose MUC, a CL-specific unlearning framework whose core innovation is Alignment Calibration—a principled, auditable optimization objective leveraging CL’s alignment property. MUC achieves controllable, representation-level forgetting after sensitive data removal via contrastive loss recalibration and embedding-space alignment constraints. The framework is architecture-agnostic, supporting mainstream CL models including SimCLR, MoCo, and CLIP, and introduces novel audit metrics enabling black-box verification and visual interpretability. Extensive experiments demonstrate that MUC consistently approaches full retraining performance across multiple benchmarks, substantially outperforming prior unlearning methods. Notably, MUC establishes the first verifiable and interpretable data unlearning capability for CL models.
This work addresses a critical limitation in existing reconstruction-based methods for learning with noisy labels: their tendency to jointly assess the reliability of observed labels and pseudo-targets, which often leads to unreliable signals substituting one another and hinders effective denoising. To overcome this, the paper proposes TRACE, a novel framework that decouples the reliability evaluation of these two sources for the first time. Specifically, it evaluates observed labels through loss fitting, shallow-feature relational stability, and prediction consistency, while assessing pseudo-targets via model confidence. These independent reliability estimates are then used to separately govern label correction and sample reweighting. By preventing error propagation, TRACE generates more trustworthy pseudo-supervision, significantly outperforming current reconstruction-based approaches across multiple synthetic and real-world noisy benchmarks, and thereby enhancing model robustness and generalization.
This work proposes a general-purpose loss function based on a Bayesian latent-variable switching mixture model that simultaneously enables robust learning and unsupervised identification of corrupted labels during training. While existing robust loss functions mitigate the adverse effects of label noise, they lack the ability to explicitly detect contaminated samples. The proposed approach uniquely unifies robust loss with interpretable posterior probabilities of label corruption, incorporates input-dependent priors to capture the spatial locality of noise, and naturally induces an Occam’s razor regularization via marginal likelihood to prevent over-flagging. Experiments on CIFAR-10 under asymmetric label noise (corruption rates 0.2–0.6) demonstrate that the method accurately recovers the underlying corruption structure, effectively distinguishes clean from noisy samples, identifies the direction of label flips, and significantly outperforms four state-of-the-art robust loss baselines.
This work systematically investigates the impact of loss weighting strategies and output parameterizations on model performance in flow matching. Through numerical experiments on both synthetic data with controllable geometric structures and real-world images, the study disentangles their interaction effects across varying data manifold dimensions, model architectures, and dataset scales, using PSNR and FID as evaluation metrics. The analysis reveals, for the first time, how the optimal choice of loss weighting and parameterization depends critically on the intrinsic structure of the data. Building on these insights, the authors formulate practical design principles that substantially improve denoising accuracy and generation quality.