noise schedule design

Designing and tuning noise injection schedules and related training objectives to enforce identity anchoring, preserve dynamics and structure across modalities, and enable stable recovery of high-quality local detail.

noisescheduledesign

12-Month Skill Trend

Momentum and market value over time
Trending
Score
+20 in 12 mo
96
12 mo agoNow
Career
Value
+$12K in 12 mo
$42K/year
12 mo agoNow

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limited generalization and robustness of deep models on clean, corrupted, and out-of-distribution (OOD) data by proposing an interleaved noise injection training strategy. During training, noisy and clean samples are alternately presented, complemented by gradient norm stabilization to enhance both feature exploration and preservation. Theoretical analysis reveals that impulsive noise is equivalent to Jacobian regularization, while Gaussian noise corresponds to curvature penalization; together, they mitigate failure modes induced by the model’s inductive bias. The approach incurs negligible computational overhead and significantly improves the robustness of both ResNet and Vision Transformer (ViT) architectures on CIFAR-100-C, ImageNet-C, and ImageNet-R benchmarks. Moreover, it seamlessly integrates with existing data augmentation techniques, yielding complementary gains.

corruption tolerancedistribution shiftnoise injection

This work addresses a critical flaw in existing evaluations of visual-language models, where output perturbations are conflated with precise adversarial concept injection, leading to inflated estimates of attack efficacy. To resolve this, the authors propose the first two-dimensional evaluation framework that explicitly distinguishes between general “influence” and true “precise injection.” Combining a programmatic drift score with large-model-assessed four-level injection grading, the framework enables fine-grained analysis of adversarial attacks. Experiments on 6,615 samples under an L∞=16/255 constraint reveal that while 66.4% of samples exhibit output perturbations, only 0.756% achieve non-zero injection, with a complete hit rate of merely 0.030%. Notably, BLIP-2 shows no significant conceptual drift at this perturbation magnitude. The study releases its full dataset and model cache to support reproducible research.

adversarial attacksmultimodal alignmentprompt injection

Noise Consistency Regularization for Improved Subject-Driven Image Synthesis

Jun 06, 2025
YN
Yao Ni
🏛️ The Australian National University | Rutgers University | Data61 CSIRO | Mitsubishi Electric Research Laboratories

In subject-driven image synthesis, fine-tuning Stable Diffusion often suffers from an imbalance between underfitting—manifesting as insufficient identity preservation—and overfitting—leading to reduced background diversity. To address this, we propose Noise Consistency Regularization (NCR), a novel regularization framework grounded in dual consistency constraints. First, we enforce noise prediction invariance over non-subject regions to preserve background fidelity and diversity. Second, we enhance the robustness of subject-specific latent codes against multiplicative noise to improve identity stability. Our method integrates CLIP-guided feature alignment, latent-space noise consistency regularization, and a robust training scheme. Experiments demonstrate that NCR consistently outperforms DreamBooth across key metrics—including CLIP similarity, background variation, and perceptual quality—effectively alleviating the identity-background trade-off inherent in subject-driven generation.

Addresses underfitting in subject-driven image synthesisEnhances subject identity and image diversityMitigates overfitting to preserve background diversity

To address insufficient robustness in multimodal representation learning, existing methods typically rely on static or heuristic noise injection, neglecting the dynamic evolution of feature distributions. This paper proposes FANoise, an adaptive noise injection framework grounded in dual perspectives—gradient dynamics and feature distribution statistics. Its core innovation is a singular-value-adaptive mechanism that dynamically modulates noise intensity according to the spectral properties of encoder features, thereby enhancing regularization while preserving training stability. Integrated within the InfoNCE contrastive learning framework, FANoise enables data-driven noise modulation. Extensive experiments across multiple vision-language models demonstrate that FANoise significantly improves generalization performance on cross-modal retrieval and understanding tasks. The method exhibits strong cross-architectural applicability and offers theoretical interpretability through its principled, spectrum-aware design.

Addresses static noise limitations in dynamic feature distributions during trainingDevelops adaptive noise injection for robust multimodal representation learningEnhances performance across various vision-language models with theoretical framework

Latest Papers

What's happening recently
View more

This work addresses the limitations of existing training-free diffusion-based style transfer methods, such as StyleID, which employ fixed style injection strengths and struggle to balance style fidelity with content preservation. We propose a novel training-free, parameter-free dynamic scheduling mechanism that adaptively modulates both style injection intensity and ControlNet geometric guidance across decoder layers and denoising timesteps. Our analysis reveals, for the first time, that a decreasing schedule—applying stronger influence in shallower layers and earlier timesteps—consistently outperforms fixed or increasing strategies, and that the two scheduling dimensions are nearly orthogonal, enabling their joint use to substantially expand the Pareto frontier. The approach demonstrates consistent gains across the Stable Diffusion backbone, achieving a state-of-the-art ArtFID of 27.036—a 6.1% improvement over StyleID—and shows comprehensive superiority across over 28,000 images and four evaluation metrics.

diffusion modelsPareto frontierstyle transfer

This work investigates the impact of parameter noise injection in stochastic gradient descent on optimization and generalization, emphasizing the need for efficient per-sample perturbations and sophisticated noise schemes. By leveraging distributional identities of linear layers, the authors propose a method that enables per-sample noise injection within mini-batches without disrupting batched computation. They systematically compare isotropic and diagonal Gaussian noise variants, demonstrating that on CIFAR-100, a lightweight single-sample isotropic Gaussian perturbation recovers most of the optimization and generalization benefits achieved by more complex multi-sample strategies. These findings suggest that simplified noise injection designs can be sufficiently effective, offering a practical alternative to computationally heavier approaches while maintaining performance gains.

generalizationmini-batch trainingnoise parameterization

This work addresses the challenge of catastrophic forgetting of source-domain identity knowledge when fine-tuning generative models under extreme few-shot conditions (fewer than 10 images), which often leads to degraded generation quality. To mitigate this issue, the authors propose a novel approach that integrates identity injection with feature consistency alignment. The method employs an identity injection module, an identity replacement mechanism, style-content disentanglement, and reconstruction modulation to effectively preserve source-domain identity information during target-domain adaptation. Extensive experiments on multiple public datasets demonstrate that the proposed method significantly outperforms current state-of-the-art techniques, achieving consistent improvements across five evaluation metrics and simultaneously enhancing both identity fidelity and overall visual quality of the generated images.

Few-Shot Generative Model AdaptationIdentity PreservationLimited Data

This work identifies a key source of object hallucination in multimodal large language models: during generation, deep-layer attention mechanisms often drift from authentic visual inputs and regress toward noise introduced in early layers. The study reveals for the first time that this phenomenon stems from such regression to early-layer noise and demonstrates that visual anchors captured in intermediate layers are crucial for reliable generation. Building on this insight, the authors propose Cross-Layer Visual Anchoring (CLVA), a training-free method that enhances features from critical intermediate layers while suppressing noise propagation from earlier stages, thereby steering attention toward accurate visual regions. CLVA consistently mitigates hallucinations across diverse model architectures and benchmarks, achieving superior performance without additional computational or memory overhead.

multimodal large language modelsobject hallucinationvisual anchors

Existing structured 3D latent diffusion models struggle with inpainting tasks due to the high sensitivity of initial noise to geometry, often failing to ensure strict alignment with surrounding context. This work proposes a training-free 3D controllable inpainting method that, for the first time, treats initial noise optimization as a control dimension independent of the sampling trajectory. By combining a backward propagation approximation derived from rectified flow models with a spectral parameterization strategy specifically designed for 3D latent variables, the method efficiently optimizes the initial noise. It achieves high-fidelity completion while significantly outperforming existing training-free baselines, demonstrating marked improvements in contextual consistency and alignment with text prompts.

3D inpaintingcontextual consistencydiffusion model

Hot Scholars

TV

Toon van Waterschoot

Professor, KU Leuven
audio processingspeech processingroom acousticsaudio engineering
MB

Michael Beigl

Professor for Informatics, Karlsruhe Institute of Technology (KIT)
Ubiquitous ComputingWearable ComputingHealth & Activity Recognition using AIInternet of Things
HY

Han Yin

Tongyi Speech Lab, Alibaba Group
Audio UnderstandingMultimodal LLM
YT

Yu Tsao

Research Fellow (Professor), Deputy Director, CITI, Academia Sinica
Assistive Oral Communication TechnologiesSpeech EnhancementVoice ConversionSpeech Assessment