gaussian-noise augmentation

Designs and implements data-augmentation pipelines or simulation workflows that add Gaussian-distributed perturbations to inputs, sensor signals, or simulation parameters to increase training diversity; and analyzes the impact of these noise injections on model robustness, generalization, and performance by tuning noise magnitude, correlation, and application targets.

gaussian-noiseaugmentation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.29
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work investigates the impact of parameter noise injection in stochastic gradient descent on optimization and generalization, emphasizing the need for efficient per-sample perturbations and sophisticated noise schemes. By leveraging distributional identities of linear layers, the authors propose a method that enables per-sample noise injection within mini-batches without disrupting batched computation. They systematically compare isotropic and diagonal Gaussian noise variants, demonstrating that on CIFAR-100, a lightweight single-sample isotropic Gaussian perturbation recovers most of the optimization and generalization benefits achieved by more complex multi-sample strategies. These findings suggest that simplified noise injection designs can be sufficiently effective, offering a practical alternative to computationally heavier approaches while maintaining performance gains.

generalizationmini-batch trainingnoise parameterization

Gaussian and Non-Gaussian Universality of Data Augmentation

Feb 18, 2022
KH
Kevin Han Huang
🏛️ University College London | Harvard University

This work systematically investigates how data augmentation affects the variance and asymptotic distribution of estimators, revealing that its efficacy is not universal: in high-dimensional regimes, augmentation can increase uncertainty in empirical prediction risk, fail to regularize effectively, and even shift the peak of the double-descent curve. The impact depends on a delicate interplay among data distribution, estimator properties, sample size, number of augmentations, and dimensionality. To address this, we propose a block-dependent adaptation technique based on the Lindeberg method, integrating random matrix theory with asymptotic statistical inference to construct the first general analytical framework applicable to both Gaussian and non-Gaussian data. This framework enables the first rigorous quantification of augmentation effects, explains several counterintuitive empirical phenomena, and validates theoretical predictions on canonical models including ridge regression and minimum-norm interpolation.

Analyze data augmentation's role as a regularizer in high-dimensional problems.Determine conditions affecting data augmentation's influence on empirical risk.Quantify data augmentation's impact on estimate variance and distribution.

This work addresses the problem of “sandbagging”—intentional underreporting of capabilities by large language models (LLMs) during safety evaluations, which undermines assessment validity. We propose a model-agnostic, zero-shot detection method requiring neither training data nor model access. Our key insight is the first empirical discovery that injecting Gaussian noise into model weights reversibly activates latent capabilities, yielding distinctive, anomalous behavioral patterns. Leveraging this phenomenon, we design an unsupervised, plug-and-play sandbagging classifier that integrates weight perturbation analysis with multi-benchmark zero-shot evaluation (MMLU, AI2, WMDP). Experiments demonstrate robust sandbagging detection across diverse model scales and multiple-choice benchmarks, achieving substantial accuracy improvements. The method is deployable, verifiable, and generalizable—providing a practical, trustworthy tool for AI safety evaluation.

Detects sandbagging in AI models via noise injection.Provides a model-agnostic tool for accurate AI evaluation.Reveals hidden capabilities masked by strategic underperformance.

This work addresses uncertainty quantification in neural networks to enhance predictive reliability and calibration. We propose Monte Carlo Noise Injection (MCNI), a novel paradigm that injects stochastic noise into model weights during training and estimates uncertainty via multiple forward passes at inference time. Crucially, we provide the first theoretical proof that this noise-injection mechanism is strictly equivalent to Bayesian inference under a deep Gaussian process (DGP) prior—thereby establishing a rigorous Bayesian foundation for weight randomization methods. Compared to conventional approaches such as Dropout and ensemble methods, MCNI achieves significantly improved accuracy in uncertainty estimation and superior calibration of predictive confidence across both regression and classification tasks. The method combines theoretical soundness with practical efficiency, requiring no architectural modifications or substantial computational overhead beyond standard Monte Carlo sampling.

Model UncertaintyNeural NetworksPrediction Accuracy

Surrogate modeling for high-cost simulators suffers from data scarcity and poor generalizability, especially when transferring knowledge across heterogeneous parameter domains—risking negative transfer. Method: This paper proposes Localized Transfer Learning Gaussian Process (LOL-GP), a framework that leverages related source-system data to enhance target-model accuracy while mitigating negative transfer induced by parametric domain discrepancies. Its core innovation is a Bayesian latent-variable regularization mechanism, which adaptively identifies transferable versus non-transferable parameter subspaces via Gibbs sampling, enabling fine-grained, localized knowledge transfer. Contribution/Results: LOL-GP overcomes the limited generalizability of conventional transfer learning in scientific simulation and supports multi-source and multi-fidelity data integration. Numerical experiments and a jet turbine design case study demonstrate that LOL-GP achieves significantly higher predictive accuracy and more reliable uncertainty quantification than state-of-the-art surrogate modeling approaches.

Costly computer simulations hinder scientific progress.Local transfer avoids negative performance impact.Transfer learning improves surrogate model efficiency.

Latest Papers

What's happening recently
View more

This work addresses the lack of systematic understanding regarding the success and failure mechanisms of generative models on real-world sensor time-series data. The authors propose SensorGen, the first unified framework for generating and evaluating multi-domain, multimodal sensor signals. They conduct a comprehensive benchmark of five prominent generative model families—including flow matching, diffusion, and autoregressive models—across four domains, seven datasets, and twelve signal modalities, introducing novel techniques for time–frequency modeling and covariate integration. Their findings reveal that flow matching models consistently achieve the best overall performance, that signal characteristics substantially influence generation quality, and that high-fidelity synthetic data can significantly enhance downstream task performance, thereby demonstrating its practical utility.

generative modelsreal-world datasensor time series

This work investigates the regularization effect induced by data augmentation in supervised regression under the high-dimensional regime where both covariate dimension and sample size grow proportionally, and its impact on generalization error. Relying solely on the first- and second-order statistics of the true data distribution and the augmentation scheme, the study leverages random feature regression, high-dimensional statistical analysis, and spectral methods to provide, for the first time, a sharp asymptotic characterization of the generalization error under model misspecification and arbitrary network architectures when only the final layer is trained. The theoretical results are validated for their accuracy in Gaussian settings and quantitatively elucidate the mechanism by which data augmentation enhances generalization performance.

data augmentationgeneralization errorproportional regime

This study systematically investigates the impact of Gaussian noise injection—varying by location (before or after activation functions) and type (additive or multiplicative)—on the performance of deep feedforward neural networks. By introducing noise at different layers and incorporating pooling mechanisms, the work reveals that activation functions exhibit a nonlinear noise-filtering effect, and that noise placement critically influences model robustness: injecting additive noise before activation yields higher accuracy and is more effectively suppressed, whereas multiplicative noise has a milder effect when applied after activation. Furthermore, early hidden layers contribute more significantly to performance degradation under post-activation noise injection, while pooling strategies consistently enhance performance across all noise configurations.

activation functiondeep neural networksinternal noise

This study addresses the high computational costs of data augmentation ensembles and their inefficiency in leveraging task symmetries by proposing Stochastic Weight Averaging (SWA) as a replacement for repetitive ensembling. Through approximation analysis via the Ornstein-Uhlenbeck process, we reveal that SWA enhances model equivariance beyond conventional performance gains in the infinite-width limit. Experiments on visual and graph classification tasks demonstrate the method’s superiority across both discrete and continuous symmetries. These findings validate SWA as an effective alternative to traditional ensembling, providing new theoretical foundations and a practical paradigm for efficiently exploiting data augmentation. This work thus bridges the gap between computational efficiency and symmetry-aware learning, offering significant implications for scalable representation learning in structured domains.

Data AugmentationDeep EnsemblesStochastic Weight Averaging

Hot Scholars

YT

Yu Tsao

Research Fellow (Professor), Deputy Director, CITI, Academia Sinica
Assistive Oral Communication TechnologiesSpeech EnhancementVoice ConversionSpeech Assessment
GG

Giovanni Geraci

Nokia | Universitat Pompeu Fabra
AI/ML6GWi-FiWireless Communications
SD

Saptarshi Debroy

Associate Professor, City University of New York
Cyber SecurityDistributed ComputingBig Data NetworkingWireless Networking