Score
Design synthetic corruptions: design and implement parameterized corruption functions and probabilistic corruption models that generate controlled perturbations (e.g., noise, blur, occlusion, compression artifacts, style or distribution shifts) to apply to inputs; build pipelines and datasets of corrupted samples and analyze how these corruptions affect algorithm performance, calibration, and robustness metrics.
Image degradation is pervasive throughout the imaging pipeline, yet existing research lacks a unified taxonomy and evaluation protocol, hindering cross-dataset and cross-task comparisons. This work introduces a causal perspective to address this gap, proposing a dual-axis classification framework: one axis categorizes degradations by their dominant causal source in the imaging pipeline—encompassing environment, sensor/optics, ISP/codec, and transmission systems—while the other characterizes their perceptual effects, augmented with a lightweight severity quantification layer. Built upon this framework, the COCO Degradation benchmark leverages PSNR, SSIM, and LPIPS to uniformly measure degradation intensity across physical artifacts, algorithmic perturbations, and perceptual distortions, substantially enhancing the evaluation of object detection model robustness under diverse imaging conditions.
Real-world distribution shifts—such as weather and illumination variations—severely degrade the robustness of deep learning models. However, collecting diverse, real-world degraded data is prohibitively expensive, prompting widespread reliance on synthetic degradation; yet its fidelity in reflecting real-world degradation effects remains unclear. Method: We construct the largest cross-domain (real vs. synthetic) semantic segmentation corruption benchmark to date, built upon Cityscapes and other datasets using the CorruptIO toolkit. We systematically evaluate 12 corruption types across multiple models and metrics (mIoU, RankCorr). Contribution/Results: We discover, for the first time, a strong correlation (ρ = 0.89) between model performance under real and synthetic corruptions. We further propose a corruption-type-level correlation analysis framework to characterize the applicability boundaries of synthetic degradation. All evaluation code, protocols, and benchmarks are publicly released to advance standardized robustness assessment.
To address the insufficient robustness of deep vision models under common image corruptions, this paper proposes a data augmentation pipeline integrating neural style transfer with controllable synthetic image generation. We first observe that stylized degradation—though increasing Fréchet Inception Distance (FID)—significantly improves corruption robustness. We further uncover the complementary mechanisms between style transfer and synthetic data augmentation, and formally characterize their compatibility boundary with rule-based methods such as TrivialAugment. Through systematic hyperparameter analysis and cross-benchmark evaluation, our method achieves state-of-the-art robust accuracy on CIFAR-10-C (93.54%), CIFAR-100-C (74.90%), and TinyImageNet-C (50.86%), establishing new SOTA results on small-scale corruption benchmarks.
Real-world structured data often suffer from demographic missingness, biased labels, and systematic sampling bias—yet existing robustness evaluations rely on random or simplistic corruptions, failing to expose worst-case vulnerabilities of high-risk ML systems. Method: We propose SAVAGE, the first causality-driven, black-box interpretable stress-testing framework for structured data. It models data dependencies via causal graphs and implements corruption templates to enable causal representation of structured data contamination. Its novel bilevel optimization algorithm supports end-to-end, targeted vulnerability discovery—even for pipelines containing non-differentiable components. Results: Experiments show that just 5% contamination generated by SAVAGE induces catastrophic performance drops, significantly outperforming baselines. Moreover, SAVAGE reveals that core assumptions underlying mainstream data cleaning and fairness-aware learning methods systematically fail under realistic data defects.
Existing data corruption studies are fragmented across specific scenarios, lacking a unified theoretical framework and systematic mitigation strategies. Method: We propose the first general corruption modeling framework based on Markov kernels, formalizing corruption as arbitrary modifications to the data distribution, hypothesis class, or loss function. We establish a provably complete taxonomy—distinguishing, for the first time, label corruption (affecting only the loss) from attribute or joint corruption (simultaneously affecting both the hypothesis class and the loss). Building on this, we introduce a generalized loss correction paradigm, deriving provably effective correction formulas for attribute and joint corruption under weaker assumptions than conventional approaches. Contribution/Results: Our framework unifies disparate corruption models and terminologies, providing a rigorous foundation for robustness analysis and algorithm design in supervised learning. It enables principled treatment of previously isolated corruption types and advances theoretical understanding of learning under distributional and structural perturbations.
This study addresses the vulnerability of mathematical agents to silently corrupted tool feedback, which impairs error detection and correction and degrades reliability. To investigate this, we propose a controlled corruption evaluation framework that employs hidden interceptors to simulate tool failures, systematically comparing the robustness of four strategies: no verification, forced in-context reflection, optional new-context verification, and structured verification. Our analysis reveals that verifier availability and verification strategy constitute independent components, demonstrating that forced reflection can fully recover performance, whereas optional verification relies on the model's proactive invocation. Experimental results indicate that without verification, accuracy drops to 72.4%, while both forced reflection and explicit detection followed by restart restore the resolution rate to 100%.
This work addresses a critical yet overlooked issue in adaptive data cleaning: fluctuating sample removal counts caused by varying partition granularities introduce budget confounding bias, leading to spurious performance gains falsely attributed to contamination identification capability. To enable fair evaluation, the authors propose an operating-point-matched assessment framework that aligns removal budgets with recall rates and incorporates threshold-agnostic metrics (AUROC and AUPRC). They systematically uncover and resolve this budget confounding problem for the first time, introducing a multi-cue adaptive cleaner—integrating learning difficulty reweighting, Euclidean distance guidance, and fine-grained partitioning—and a false positive decomposition analysis. Experiments on CIFAR-10 and ImageNet-100 reveal that most existing methods lose their apparent advantage under matched operating points, demonstrating genuine efficacy only under low contamination rates or high-recall, heavily corrupted scenarios.
This work addresses the degradation of coverage guarantees in online conformal prediction when feedback is corrupted by noise, communication failures, or adversarial attacks. It presents the first systematic modeling of arbitrary binary feedback corruption and provides explicit miscoverage bounds under two distinct error models: independent random flips and memory-bounded adversarial errors. To mitigate the impact of corrupted feedback, the paper introduces two robust mechanisms—a threshold-based feedback filtering scheme and an active compensation strategy—effectively preserving predictive reliability. The proposed approach integrates online conformal prediction with robust statistical learning, making it suitable for non-stationary sequential environments. Empirical evaluations on real-world datasets demonstrate substantial improvements over existing baselines, achieving better calibration and significantly smaller prediction sets.
为解决多模态医学图像分割中因质量差异导致的融合失败问题,提出CoReFuse-Med框架,通过抑制特征传输中的损坏并重新平衡模态贡献来提高准确性和鲁棒性。
This study addresses the pervasive issues of semantic mislabeling and bounding box localization errors in object detection datasets by systematically evaluating, for the first time, the effectiveness of training-free feature-space methods for annotation error detection. Leveraging multiple pretrained embedding models, the approach is rigorously tested on both synthetic noise—including symmetric, asymmetric, and localization-type perturbations—and real-world annotation errors in the VOC2012 and KITTI datasets. Experimental results demonstrate that feature-space methods are highly effective at identifying semantic mislabels but exhibit limited capability in detecting localization inaccuracies. To facilitate future research, the authors publicly release all code and a curated set of verified erroneous annotations, establishing a valuable benchmark for the community.
This study investigates the robust learning mechanisms of neural networks when input data are severely corrupted by attribute noise—such as additive or replacement noise—while labels remain intact. Through experiments with multilayer perceptrons, mean-field analysis of infinite-width networks, and prototype-based classification theory, the work reveals for the first time that networks consistently adopt a nearest-class-mean decision rule even when over 90% of the input features are corrupted. This behavior is shown to be universal across network depth, activation functions, and noise distributions. Empirical results demonstrate that finite-width networks closely align with theoretical predictions, significantly outperforming random guessing under extreme noise conditions. These findings establish an interpretable and analytically tractable foundation for understanding robustness in deep learning.