Score
Designs and implements controlled synthetic-noise and corruption generators and injection pipelines that add specified distributions and magnitudes of perturbation — e.g., additive/multiplicative noise, heavy-tailed outliers, blur models, label or syntactic perturbations, or diff-style (diffusion-style) disturbances — to data or labels while parameterizing severity and preserving required semantics. Builds experiments and evaluation protocols that use these injections to test and benchmark model robustness, study privacy/utility trade-offs, simulate real-world corruptions, and tune injection parameters for reproducible comparisons.
Image degradation is pervasive throughout the imaging pipeline, yet existing research lacks a unified taxonomy and evaluation protocol, hindering cross-dataset and cross-task comparisons. This work introduces a causal perspective to address this gap, proposing a dual-axis classification framework: one axis categorizes degradations by their dominant causal source in the imaging pipeline—encompassing environment, sensor/optics, ISP/codec, and transmission systems—while the other characterizes their perceptual effects, augmented with a lightweight severity quantification layer. Built upon this framework, the COCO Degradation benchmark leverages PSNR, SSIM, and LPIPS to uniformly measure degradation intensity across physical artifacts, algorithmic perturbations, and perceptual distortions, substantially enhancing the evaluation of object detection model robustness under diverse imaging conditions.
This study addresses a critical gap in synthetic data generation (SDG) research, which has predominantly focused on privacy attacks initiated by data recipients while overlooking internal adversaries—such as data owners or generators—who may degrade data quality by perturbing real data. The work formally introduces this internal threat model and proposes targeted perturbation strategies based on label flipping and feature importance manipulation. Through systematic evaluation across multiple mainstream SDG frameworks, the experiments demonstrate that even minimal perturbations can substantially impair downstream task performance and amplify statistical distributional biases. These findings reveal a pronounced vulnerability in current SDG pipelines regarding data integrity and underscore the urgent need for robustness and integrity verification mechanisms in synthetic data workflows.
Diffusion models suffer from weak controllability during sampling, making it difficult to satisfy statistical constraints—primarily because the relationship between initial noise perturbations and final outputs remains uncharacterized. This work provides the first theoretical proof that, under diffusion ODE sampling, the output exhibits a strongly linear response to initial noise perturbations. Leveraging this insight, we propose CCS (Controlled Controllable Sampling), the first explicit noise-space controllable sampling framework. CCS employs differentiable controllers in the noise space to precisely specify target statistics—e.g., mean and variance—without altering the model architecture or requiring retraining. Evaluated across multiple benchmarks, CCS achieves state-of-the-art statistical controllability while preserving high sample quality (FID ≤ 2.1) and diversity (LPIPS ≥ 0.43).
Real-world structured data often suffer from demographic missingness, biased labels, and systematic sampling bias—yet existing robustness evaluations rely on random or simplistic corruptions, failing to expose worst-case vulnerabilities of high-risk ML systems. Method: We propose SAVAGE, the first causality-driven, black-box interpretable stress-testing framework for structured data. It models data dependencies via causal graphs and implements corruption templates to enable causal representation of structured data contamination. Its novel bilevel optimization algorithm supports end-to-end, targeted vulnerability discovery—even for pipelines containing non-differentiable components. Results: Experiments show that just 5% contamination generated by SAVAGE induces catastrophic performance drops, significantly outperforming baselines. Moreover, SAVAGE reveals that core assumptions underlying mainstream data cleaning and fairness-aware learning methods systematically fail under realistic data defects.
To address slow convergence, poor distribution alignment, and the privacy–utility trade-off in generative modeling of high-dimensional privacy-sensitive data (e.g., biomedical datasets), this paper proposes a differential privacy (DP) synthetic data generation framework based on latent-space noise injection. Methodologically, it introduces the first local (ε,δ)-DP perturbation directly into the latent variable layer of Masked Autoregressive Flows (MAF), coupled with an invertible mapping to ensure bijective correspondence between original and synthetic data. A single tunable parameter governs the privacy budget, and √n statistical consistency is recovered under meta-analytic aggregation. Experiments demonstrate a significant reduction in Wasserstein distance, membership inference attack success rates below 5%, and asymptotic efficiency of aggregated estimators matching classical statistical benchmarks.
This work addresses the problem of “sandbagging”—intentional underreporting of capabilities by large language models (LLMs) during safety evaluations, which undermines assessment validity. We propose a model-agnostic, zero-shot detection method requiring neither training data nor model access. Our key insight is the first empirical discovery that injecting Gaussian noise into model weights reversibly activates latent capabilities, yielding distinctive, anomalous behavioral patterns. Leveraging this phenomenon, we design an unsupervised, plug-and-play sandbagging classifier that integrates weight perturbation analysis with multi-benchmark zero-shot evaluation (MMLU, AI2, WMDP). Experiments demonstrate robust sandbagging detection across diverse model scales and multiple-choice benchmarks, achieving substantial accuracy improvements. The method is deployable, verifiable, and generalizable—providing a practical, trustworthy tool for AI safety evaluation.
This work investigates the impact of parameter noise injection in stochastic gradient descent on optimization and generalization, emphasizing the need for efficient per-sample perturbations and sophisticated noise schemes. By leveraging distributional identities of linear layers, the authors propose a method that enables per-sample noise injection within mini-batches without disrupting batched computation. They systematically compare isotropic and diagonal Gaussian noise variants, demonstrating that on CIFAR-100, a lightweight single-sample isotropic Gaussian perturbation recovers most of the optimization and generalization benefits achieved by more complex multi-sample strategies. These findings suggest that simplified noise injection designs can be sufficiently effective, offering a practical alternative to computationally heavier approaches while maintaining performance gains.
Existing open-source tools lack a unified infrastructure, making it challenging to efficiently support the diverse modeling choices and privacy requirements inherent in synthetic data generation. This work proposes tidysynthesis, a modular and extensible meta-package that integrates a wide range of statistical modeling and differential privacy techniques through a unified declarative API. For the first time, it enables cross-framework algorithm composition and flexible customization of synthetic data workflows. The system substantially enhances both development efficiency and privacy guarantees, with its completeness, usability, and capacity to handle complex synthesis tasks demonstrated through empirical validation on U.S. Census survey data.
This work addresses the triple challenges of controllability, content fidelity, and safety in diffusion-based image editing. It systematically investigates paradigms including text/mask guidance, point-and-drag manipulation, and inversion mapping, formalizing editing objectives and analyzing the dynamics of noise injection, score guidance, and inversion errors. The authors propose a unified framework integrating mask localization and instruction-guided editing, establishing— for the first time—theoretical bounds on reconstruction error, stability under repeated edits, and locality of modifications. Their analysis reveals failure modes in existing methods, such as identity drift, prompt sensitivity, and compositional errors. Through multidimensional evaluation using FID, identity similarity, CLIP alignment, and artifact scoring, the study comprehensively characterizes the controllability–fidelity trade-off and explores concept erasure techniques like MACE and ANT as ethical safeguards.
This work addresses the limited generalization and robustness of deep models on clean, corrupted, and out-of-distribution (OOD) data by proposing an interleaved noise injection training strategy. During training, noisy and clean samples are alternately presented, complemented by gradient norm stabilization to enhance both feature exploration and preservation. Theoretical analysis reveals that impulsive noise is equivalent to Jacobian regularization, while Gaussian noise corresponds to curvature penalization; together, they mitigate failure modes induced by the model’s inductive bias. The approach incurs negligible computational overhead and significantly improves the robustness of both ResNet and Vision Transformer (ViT) architectures on CIFAR-100-C, ImageNet-C, and ImageNet-R benchmarks. Moreover, it seamlessly integrates with existing data augmentation techniques, yielding complementary gains.