Score
Designs and implements training curricula and loss formulations that mix clean and corrupted inputs, control the proportion and difficulty of corruptions over time (e.g., gradually increasing corruption severity), and enforce balanced objectives between clean and corrupted views. These artifacts are built and analyzed to train models that retain performance across diverse corruption or failure patterns while avoiding bias toward either clean or corrupted data.
Existing data corruption studies are fragmented across specific scenarios, lacking a unified theoretical framework and systematic mitigation strategies. Method: We propose the first general corruption modeling framework based on Markov kernels, formalizing corruption as arbitrary modifications to the data distribution, hypothesis class, or loss function. We establish a provably complete taxonomy—distinguishing, for the first time, label corruption (affecting only the loss) from attribute or joint corruption (simultaneously affecting both the hypothesis class and the loss). Building on this, we introduce a generalized loss correction paradigm, deriving provably effective correction formulas for attribute and joint corruption under weaker assumptions than conventional approaches. Contribution/Results: Our framework unifies disparate corruption models and terminologies, providing a rigorous foundation for robustness analysis and algorithm design in supervised learning. It enables principled treatment of previously isolated corruption types and advances theoretical understanding of learning under distributional and structural perturbations.
This work investigates the impact mechanisms of hallucination, erroneous responses, and low-quality OCR—collectively termed “contamination”—on multimodal large language models (MLLMs) during visual instruction tuning (VIT). We find that contamination-induced degradation is superficial, primarily affecting output-layer parameters; thus, freezing lower-layer parameters or fine-tuning with merely 1% clean data suffices to restore over 95% of original performance. Building on this insight, we propose the first external-label-free self-verifying data cleaning framework: it identifies contaminated samples via parameter plasticity analysis, then integrates self-supervised confidence estimation with contamination-aware lightweight post-training, forming a two-stage robust debiasing paradigm. Our method significantly outperforms existing approaches across multiple VIT benchmarks and enables end-to-end automatic data cleaning.
本文探讨了在数据不完美条件下的机器学习挑战,并通过信息损失、经验风险偏差等机制组织代表性方法,如重建生成、再平衡与表示校准等。
This work addresses the challenge of robust training on tabular data suffering from heterogeneous corruptions—such as noise, missing values, and feature bias—when only column-level reliability indicators are available. The authors propose a quality-aware training mechanism that uniquely integrates column-wise reliability priors directly into the optimization dynamics. By jointly optimizing a learnable feature modulation layer and a quality-dependent proximal regularizer, the method adaptively adjusts the contribution of features according to their trustworthiness, without requiring explicit data imputation or sample reweighting. Extensive experiments across 50 classification and regression datasets demonstrate that the proposed approach significantly outperforms existing baselines, exhibiting exceptional robustness particularly in low-data regimes and under systematic feature biases.
This paper addresses the challenge of robust federated learning under model heterogeneity and client-side data corruption—including noise and compression artifacts. We propose the first robust federated learning framework tailored for asymmetric heterogeneous settings. Methodologically, we innovatively integrate diversity-enhanced supervised contrastive learning with a selective one-way collaboration mechanism: the former strengthens robust cross-architecture feature representation, while the latter enables the server to actively reject low-quality client updates, facilitating adaptive knowledge transfer. Key technical components include hybrid data augmentation, asymmetric model aggregation, and client selection strategies. Extensive experiments under diverse data corruption and model heterogeneity scenarios demonstrate that our approach significantly improves both global model accuracy and convergence stability, consistently outperforming existing state-of-the-art methods.
Transformer models exhibit fragile generalization on long sequences and structurally complex inputs. Method: We propose an automated curriculum learning framework driven by failure cases, centered on an executable-validator-guided counterexample discovery mechanism that dynamically generates challenging instances and constructs adaptive training curricula—without manual difficulty annotation. Our approach integrates counterexample-informed data augmentation, logical verification constraints, and Transformer fine-tuning to enable continuous self-correction. Contribution/Results: Experiments demonstrate a 30× improvement in sequence-length extrapolation on algorithmic reasoning and natural language tasks. Compared to uniform data augmentation, our method achieves a 3.75× speedup in computational efficiency and significantly outperforms both static training and conventional curriculum learning baselines.
This study investigates the robust learning mechanisms of neural networks when input data are severely corrupted by attribute noise—such as additive or replacement noise—while labels remain intact. Through experiments with multilayer perceptrons, mean-field analysis of infinite-width networks, and prototype-based classification theory, the work reveals for the first time that networks consistently adopt a nearest-class-mean decision rule even when over 90% of the input features are corrupted. This behavior is shown to be universal across network depth, activation functions, and noise distributions. Empirical results demonstrate that finite-width networks closely align with theoretical predictions, significantly outperforming random guessing under extreme noise conditions. These findings establish an interpretable and analytically tractable foundation for understanding robustness in deep learning.
This work addresses the insufficient robustness of current models under natural image corruptions, particularly their vulnerability in safety-critical scenarios. The authors present the first explicit characterization of internal robust computational pathways within neural networks, revealing a consistent attenuation of robust features across layers. To counteract this degradation, they propose a novel “Suppress and Diversify” mechanism that is architecture-agnostic, parameter-free, and incurs zero overhead at test time. This approach dynamically selects and diversifies symmetry-preserving robust pathways to enhance overall model robustness. Extensive experiments across eight benchmarks demonstrate that the method consistently improves performance across diverse vision tasks, backbone architectures, and complex real-world conditions, highlighting its strong generalizability and scalability.
研究探讨了在存在单调对抗性干扰的情况下,多分类及部分二元概念类学习问题的难度显著增加,并提出当新增干扰数据量为o(n)时,所有类别仍可学习。
This study addresses the limitation of relying on a single uncertainty metric to simultaneously detect heterogeneous data corruptions, such as image noise and label flipping, in federated learning. Employing a ResNet-20 architecture with Monte Carlo Dropout under Dirichlet non-IID data partitioning, this work compares the detection performance—evaluated via AUC—of predictive loss and entropy-based uncertainty. It proposes an adaptive strategy that matches detection signals to specific corruption types. The findings reveal a complementary mechanism: predictive loss is more sensitive to label flipping (AUC > 0.85), whereas entropy uncertainty proves superior for detecting image noise. By overcoming the constraints of single-metric uncertainty estimation, this research establishes an adaptive data quality assessment framework tailored for federated learning environments.