Score
Designs, implements, or evaluates normalization modules that condition the computation of feature statistics (means, variances) and affine parameters (scale, bias) on external or process-specific signals so network activations, embeddings, patches, or vectors are modulated by context or mode. This includes conditional and dual‑mode variants of layer/embedding/patch normalization that enable mode‑specific adaptation, alignment between representations, or preservation of semantic features across operating conditions.
This work investigates how the placement of normalization layers (e.g., BatchNorm, LayerNorm) within neural network hidden layers affects the predictive distribution at initialization—particularly the initial bias across class predictions. Method: Leveraging a novel synthesis of random matrix theory and Gaussian process approximations, we establish the first theoretical link between normalization position and initialization-time prediction statistics, introducing the “initialization bias degree” as a quantitative metric. Results: We demonstrate that pre-normalization systematically attenuates initial class bias, driving the output logits toward a uniform, neutral distribution at initialization. This mechanism markedly improves training stability and convergence consistency across architectures—including fully connected networks, CNNs, and Transformers—without architectural or hyperparameter modification. Empirical validation on multiple benchmarks confirms its interpretable, path-level control over optimization dynamics.
This work theoretically characterizes multi-scale feature learning in neural networks, unifying the understanding of kernel scaling and data-driven kernel adaptation—and their respective expressive power and inductive biases—across training regimes from the mean-field to the standard (NTK) scaling limit. Method: We propose the first unified analytical framework spanning all scaling regimes; prove, for the first time, that under mean prediction, kernel adaptation is equivalent to effective kernel scaling while preserving directional feature learning; and derive a statistically exact closed-form expression for network outputs across the full scaling spectrum via statistical mechanics, continuum approximation, and higher-order saddle-point analysis. Contribution/Results: Our core contribution is a rigorous theory of multi-scale adaptive feature learning that bridges long-standing conceptual divides between scaling paradigms, revealing the precise mechanisms of feature emergence and fundamental limits of expressivity—thereby providing a unified foundation for understanding deep learning principles.
The optimal distribution of hidden-layer activations in deep neural networks lacks theoretical grounding and practical implementation strategies. This paper systematically establishes the information-theoretic advantages of feature Gaussianization and proposes Normality Normalization—a differentiable, plug-and-play layer that unifies distributional shape correction (via Box–Cox or Yeo–Johnson power transformations) with robustness enhancement (via additive Gaussian noise during training and moment-based normalization). The method explicitly steers hidden activations toward a standard normal distribution. Empirically, it significantly improves generalization across diverse architectures and datasets; exhibits robustness to variations in network width, depth, and batch size; and enhances resilience against both adversarial perturbations and random input noise.
Computing approximate vanishing ideals from scarce, uncertain data points suffers from poor numerical stability under traditional coefficient-norm-based polynomial normalization, which is scale-sensitive and lacks robustness. Method: We propose a gradient-weighted, data-driven seminorm normalization strategy—novelly integrated into border basis construction—to replace conventional normalization. This approach rigorously guarantees scale invariance and significantly enhances robustness against data perturbations. Contribution/Results: Theoretical analysis and experiments on three affine varieties demonstrate that the new method completely eliminates scale dependence, improves noise resilience substantially, and requires only minor algorithmic adjustments without increasing time complexity. Its core contribution is a geometrically meaningful and numerically stable normalization framework, establishing a new paradigm for approximate algebraic geometry modeling.
In multimodal learning, modality imbalance often stems from internal optimization bias within neural networks—not merely from inherent representational disparities across modalities. Existing approaches suppress dominant modalities to boost underperforming ones, inadvertently degrading overall performance. To address this, we propose Adaptive Network-Intrinsic Modulation (ANIM), the first method to decouple suboptimally trained parameters in dominant modalities and introduce lightweight auxiliary modules. ANIM jointly analyzes parameter-level optimization states and cross-layer modality imbalance, dynamically adjusting modulation strength across network depths to enable synergistic optimization of both dominant and underperforming modalities. Importantly, ANIM is architecture-agnostic—requiring no modifications to backbone networks, fusion strategies, or optimizers—ensuring broad applicability. Extensive experiments on multiple benchmarks demonstrate significant improvements over state-of-the-art methods, achieving superior modality balance while simultaneously enhancing overall model performance.
This work addresses the limited interpretability of existing parameter-efficient fine-tuning methods at the neuron level, which obscures how models reuse or bypass internal computations. Inspired by neuromodulation, the authors propose an activation-space fine-tuning approach that models adaptation as a “mode-switching” mechanism. By freezing the backbone weights and learning per-layer trainable thresholds and gains for neurons—combined with smooth gating during training and a hardening strategy at inference—the method enables conditional computation and neuron-level attribution. Evaluated on the rotated MNIST task, it introduces only a few hundred parameters per layer, significantly outperforming frozen baselines while achieving partial activation sparsity and offering clear interpretability of individual neuron activations.
Existing interpretability methods struggle to distinguish whether model components genuinely encode a target capability or merely propagate upstream signals. This work proposes Weight Patching, a source-directed intervention in weight space that operates on isomorphic models exhibiting varying behavioral strengths. By substituting specific module weights and anchoring behavioral interfaces via vector alignment, the method precisely localizes source-level mechanisms within large language models. The framework enables, for the first time, tracing the pathway of capability transmission from shallow source carriers to downstream execution circuits, thereby supporting mechanism-aware model merging. Experiments on instruction-following tasks successfully identify critical mechanistic components, significantly improving selective fusion of expert models, with findings further validated externally.
This work investigates the mechanism of lightweight yet critical learnable scaling vectors in large language models. Through theoretical analysis and large-scale pretraining experiments, it reveals for the first time that, within Pre-Norm architectures, these vectors primarily improve optimization dynamics via a self-amplifying preconditioning effect rather than enhancing model expressivity. The study further identifies divergent responses to weight decay between Input-Norm and Output-Norm layers. Building on these insights, three efficient improvements are proposed: heterogeneous branch design, optimized placement of linear mappings, and a magnitude-direction decoupled reparameterization. Evaluated across dense and mixture-of-experts models ranging from 0.12B to 2B parameters, the unified scaling strategy consistently reduces final training loss, enhances stability, and incurs negligible additional parameters or computational cost.
This study investigates the impact of normalization strategies in time series preprocessing on the representational capacity of Transformer models. Focusing on commonly used methods such as Standard and Min-Max normalization, it provides the first theoretical analysis of how these techniques influence the discriminative power of the representation space and introduces a quantitative evaluation framework to assess this capability. Through theoretical bounds and systematic experiments across multiple benchmark datasets—complemented by comparisons between instance-wise normalization and global scaling—the work demonstrates that normalization significantly affects model performance, yet no universally optimal strategy exists. Notably, for certain tasks, omitting normalization altogether yields superior results, revealing that preprocessing choices must be co-designed with task-specific characteristics.
Function-parameterized neural networks are highly sensitive to initialization, and conventional data-agnostic initialization schemes often fail to capture the structural characteristics of target signals, leading to slow convergence and unstable performance. This work proposes a prior-guided initialization strategy that, for the first time, integrates data-driven spectral priors into both network initialization and architecture design. Specifically, fast Fourier transform (FFT) is employed to extract seasonal priors that inform model depth and initial state, while residual regression is used to parameterize trend components. Without altering the training procedure, the proposed method significantly accelerates convergence, reduces performance variance, and improves computational efficiency across both synthetic and real-world datasets. Notably, it maintains reconstruction accuracy even when using a lower-dimensional encoder, consistently outperforming standard initialization approaches.