process-conditioned normalization

Designs, implements, or evaluates normalization modules that condition the computation of feature statistics (means, variances) and affine parameters (scale, bias) on external or process-specific signals so network activations, embeddings, patches, or vectors are modulated by context or mode. This includes conditional and dual‑mode variants of layer/embedding/patch normalization that enable mode‑specific adaptation, alignment between representations, or preservation of semantic features across operating conditions.

process-conditionednormalization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.77
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$205K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Where You Place the Norm Matters: From Prejudiced to Neutral Initializations

May 16, 2025
EF
Emanuele Francazi
🏛️ EPFL | La Sapienza University of Rome | University of Basel | Eawag (ETH)

This work investigates how the placement of normalization layers (e.g., BatchNorm, LayerNorm) within neural network hidden layers affects the predictive distribution at initialization—particularly the initial bias across class predictions. Method: Leveraging a novel synthesis of random matrix theory and Gaussian process approximations, we establish the first theoretical link between normalization position and initialization-time prediction statistics, introducing the “initialization bias degree” as a quantitative metric. Results: We demonstrate that pre-normalization systematically attenuates initial class bias, driving the output logits toward a uniform, neutral distribution at initialization. This mechanism markedly improves training stability and convergence consistency across architectures—including fully connected networks, CNNs, and Transformers—without architectural or hyperparameter modification. Empirical validation on multiple benchmarks confirms its interpretable, path-level control over optimization dynamics.

How normalization placement affects initial network predictionsImpact of normalization on class prediction distribution at initializationLinking architectural choices to early training behavior dynamics

From Kernels to Features: A Multi-Scale Adaptive Theory of Feature Learning

Feb 05, 2025
NR
Noa Rubin
🏛️ The Hebrew University of Jerusalem | Jülich Research Centre | RWTH Aachen University | Tel-Aviv University

This work theoretically characterizes multi-scale feature learning in neural networks, unifying the understanding of kernel scaling and data-driven kernel adaptation—and their respective expressive power and inductive biases—across training regimes from the mean-field to the standard (NTK) scaling limit. Method: We propose the first unified analytical framework spanning all scaling regimes; prove, for the first time, that under mean prediction, kernel adaptation is equivalent to effective kernel scaling while preserving directional feature learning; and derive a statistically exact closed-form expression for network outputs across the full scaling spectrum via statistical mechanics, continuum approximation, and higher-order saddle-point analysis. Contribution/Results: Our core contribution is a rigorous theory of multi-scale adaptive feature learning that bridges long-standing conceptual divides between scaling paradigms, revealing the precise mechanisms of feature emergence and fundamental limits of expressivity—thereby providing a unified foundation for understanding deep learning principles.

Bridges multi-scale adaptive feature learning approachesDerives network output statistics across scaling regimesReduces kernel adaptation to effective kernel rescaling

On the Importance of Gaussianizing Representations

May 01, 2025
DE
Daniel Eftekhari
🏛️ University of Toronto | Vector Institute

The optimal distribution of hidden-layer activations in deep neural networks lacks theoretical grounding and practical implementation strategies. This paper systematically establishes the information-theoretic advantages of feature Gaussianization and proposes Normality Normalization—a differentiable, plug-and-play layer that unifies distributional shape correction (via Box–Cox or Yeo–Johnson power transformations) with robustness enhancement (via additive Gaussian noise during training and moment-based normalization). The method explicitly steers hidden activations toward a standard normal distribution. Empirically, it significantly improves generalization across diverse architectures and datasets; exhibits robustness to variations in network width, depth, and batch size; and enhances resilience against both adversarial perturbations and random input noise.

Addressing lack of prescribed distribution for neural network activationsEncouraging normality in feature representations using power transformImproving model robustness to random perturbations via Gaussian noise

Gradient-Weighted, Data-Driven Normalization for Approximate Border Bases -- Concept and Computation

Jun 11, 2025
HK
Hiroshi Kera
🏛️ Chiba University | Zuse Institute Berlin | Rhine-Waal University of Applied Sciences

Computing approximate vanishing ideals from scarce, uncertain data points suffers from poor numerical stability under traditional coefficient-norm-based polynomial normalization, which is scale-sensitive and lacks robustness. Method: We propose a gradient-weighted, data-driven seminorm normalization strategy—novelly integrated into border basis construction—to replace conventional normalization. This approach rigorously guarantees scale invariance and significantly enhances robustness against data perturbations. Contribution/Results: Theoretical analysis and experiments on three affine varieties demonstrate that the new method completely eliminates scale dependence, improves noise resilience substantially, and requires only minor algorithmic adjustments without increasing time complexity. Its core contribution is a geometrically meaningful and numerically stable normalization framework, establishing a new paradigm for approximate algebraic geometry modeling.

Adapting border basis concept for approximate treatment of uncertain data pointsDemonstrating superior robustness and invariance compared to coefficient normalizationProposing gradient-weighted normalization for better stability and scaling invariance

In multimodal learning, modality imbalance often stems from internal optimization bias within neural networks—not merely from inherent representational disparities across modalities. Existing approaches suppress dominant modalities to boost underperforming ones, inadvertently degrading overall performance. To address this, we propose Adaptive Network-Intrinsic Modulation (ANIM), the first method to decouple suboptimally trained parameters in dominant modalities and introduce lightweight auxiliary modules. ANIM jointly analyzes parameter-level optimization states and cross-layer modality imbalance, dynamically adjusting modulation strength across network depths to enable synergistic optimization of both dominant and underperforming modalities. Importantly, ANIM is architecture-agnostic—requiring no modifications to backbone networks, fusion strategies, or optimizers—ensuring broad applicability. Extensive experiments on multiple benchmarks demonstrate significant improvements over state-of-the-art methods, achieving superior modality balance while simultaneously enhancing overall model performance.

Adaptively modulates learning across different network depthsAddresses optimization bias within multimodal learning networksBalances modality learning without hindering dominant or weak modalities

Latest Papers

What's happening recently
View more

This work addresses the limited interpretability of existing parameter-efficient fine-tuning methods at the neuron level, which obscures how models reuse or bypass internal computations. Inspired by neuromodulation, the authors propose an activation-space fine-tuning approach that models adaptation as a “mode-switching” mechanism. By freezing the backbone weights and learning per-layer trainable thresholds and gains for neurons—combined with smooth gating during training and a hardening strategy at inference—the method enables conditional computation and neuron-level attribution. Evaluated on the rotated MNIST task, it introduces only a few hundred parameters per layer, significantly outperforming frozen baselines while achieving partial activation sparsity and offering clear interpretability of individual neuron activations.

activation sparsityconditional computationmodel interpretability

Existing interpretability methods struggle to distinguish whether model components genuinely encode a target capability or merely propagate upstream signals. This work proposes Weight Patching, a source-directed intervention in weight space that operates on isomorphic models exhibiting varying behavioral strengths. By substituting specific module weights and anchoring behavioral interfaces via vector alignment, the method precisely localizes source-level mechanisms within large language models. The framework enables, for the first time, tracing the pathway of capability transmission from shallow source carriers to downstream execution circuits, thereby supporting mechanism-aware model merging. Experiments on instruction-following tasks successfully identify critical mechanistic components, significantly improving selective fusion of expert models, with findings further validated externally.

behavioral capabilityLLMsmechanistic interpretability

This work investigates the mechanism of lightweight yet critical learnable scaling vectors in large language models. Through theoretical analysis and large-scale pretraining experiments, it reveals for the first time that, within Pre-Norm architectures, these vectors primarily improve optimization dynamics via a self-amplifying preconditioning effect rather than enhancing model expressivity. The study further identifies divergent responses to weight decay between Input-Norm and Output-Norm layers. Building on these insights, three efficient improvements are proposed: heterogeneous branch design, optimized placement of linear mappings, and a magnitude-direction decoupled reparameterization. Evaluated across dense and mixture-of-experts models ranging from 0.12B to 2B parameters, the unified scaling strategy consistently reduces final training loss, enhances stability, and incurs negligible additional parameters or computational cost.

expressivitylarge language modelsnormalization layers

This study investigates the impact of normalization strategies in time series preprocessing on the representational capacity of Transformer models. Focusing on commonly used methods such as Standard and Min-Max normalization, it provides the first theoretical analysis of how these techniques influence the discriminative power of the representation space and introduces a quantitative evaluation framework to assess this capability. Through theoretical bounds and systematic experiments across multiple benchmark datasets—complemented by comparisons between instance-wise normalization and global scaling—the work demonstrates that normalization significantly affects model performance, yet no universally optimal strategy exists. Notably, for certain tasks, omitting normalization altogether yields superior results, revealing that preprocessing choices must be co-designed with task-specific characteristics.

expressivitynormalizationscaling

Function-parameterized neural networks are highly sensitive to initialization, and conventional data-agnostic initialization schemes often fail to capture the structural characteristics of target signals, leading to slow convergence and unstable performance. This work proposes a prior-guided initialization strategy that, for the first time, integrates data-driven spectral priors into both network initialization and architecture design. Specifically, fast Fourier transform (FFT) is employed to extract seasonal priors that inform model depth and initial state, while residual regression is used to parameterize trend components. Without altering the training procedure, the proposed method significantly accelerates convergence, reduces performance variance, and improves computational efficiency across both synthetic and real-world datasets. Notably, it maintains reconstruction accuracy even when using a lower-dimensional encoder, consistently outperforming standard initialization approaches.

convergencedata priorsfunction parameterization

Hot Scholars

XQ

Xianbiao Qi

Shenzhen Intellifusion Technologies Co., Ltd.
Neural Network OptimizationGenerative ModelsLarge-Scale Pretrain ModelsOCR
NS

Nicu Sebe

University of Trento
computer visionmultimedia
XZ

Xun Zhou

Professor of Computer Science, Harbin Institute of Technology, Shenzhen (HIT-SZ)
Big data analyticsSpatial databaseSpatial Data MiningGIS
ZL

Zhouchen Lin

Professor, Peking University; Fellow of IEEE, IAPR, CSIG & AAIA; ex-VP of Samsung Research
machine learningcomputer visionimage processingnumerical optimization
CG

Chun-Guang Li

Associate Professor, Beijing University of Posts and Telecommunications
Subspace ClusteringSelf-Supervised LearningTime Series ModelingBiomedical Engineering