dataset-specific normalization

Design and learn normalization transforms that are tailored to the statistical properties of a specific dataset, implemented as parameterized operations fitted to that dataset. These transforms are often invertible so they standardize inputs for models while preserving the ability to exactly or approximately reconstruct original-scale data for downstream use and evaluation.

dataset-specificnormalization

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.37
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

A Power Transform

Feb 15, 2025
JT
Jonathan T. Barron

This paper addresses the narrow applicability and lack of unification across existing power transformations (e.g., Box–Cox, Tukey) for modeling diverse multivariate mathematical objects. We propose a generalized power transformation framework based on a differentiable family of generalized power functions, parameterized by a single continuous scalar. For the first time, this framework enables unified parametric representation of loss functions, kernel functions, probability densities, convex/concave functions, and neural network activation functions. Theoretically, we rigorously establish its mathematical equivalence and functional generalization capacity. Empirically, it outperforms conventional methods in data standardization, distribution fitting, and neural activation tasks. Our core contribution is the development of the first unified power transformation paradigm that simultaneously ensures universality, analytical tractability, and computational feasibility—providing a foundational tool for statistical modeling and machine learning.

Develops a novel power transformEnhances dataset normalization methodsUnifies diverse mathematical functions

This study investigates the impact of normalization strategies in time series preprocessing on the representational capacity of Transformer models. Focusing on commonly used methods such as Standard and Min-Max normalization, it provides the first theoretical analysis of how these techniques influence the discriminative power of the representation space and introduces a quantitative evaluation framework to assess this capability. Through theoretical bounds and systematic experiments across multiple benchmark datasets—complemented by comparisons between instance-wise normalization and global scaling—the work demonstrates that normalization significantly affects model performance, yet no universally optimal strategy exists. Notably, for certain tasks, omitting normalization altogether yields superior results, revealing that preprocessing choices must be co-designed with task-specific characteristics.

expressivitynormalizationscaling

IN-Flow: Instance Normalization Flow for Non-stationary Time Series Forecasting

Jan 30, 2024
WF
Wei Fan
🏛️ University of Oxford | Microsoft Research | University of Macau | University of Central Florida | Arizona State University

To address performance degradation in non-stationary time series forecasting caused by distributional shift, this paper proposes a decoupled modeling framework that separates distribution correction from forecasting and introduces a bilevel optimization paradigm for joint learning. Its core innovation is Instance Normalization Flow (IN-Flow)—a reversible, bidirectional, and highly expressive temporal distribution transformation network explicitly designed for forecasting, overcoming the limitation of conventional normalizing flows restricted to generative tasks. IN-Flow integrates instance normalization layers with stacked invertible neural networks: the outer level optimizes distribution transformation, while the inner level optimizes forecasting, ensuring compatibility with arbitrary forecasting architectures and eliminating reliance on statistical assumptions. Extensive experiments on synthetic and diverse real-world datasets demonstrate significant improvements over state-of-the-art methods, strong robustness to unseen distribution shifts, and simultaneous gains in both predictive accuracy and generalization capability.

Addresses non-stationarity in time series forecasting.Introduces IN-Flow for effective time series transformation.Proposes decoupled formulation for distribution shift.

Principled Interpolation in Normalizing Flows

Oct 22, 2020
SG
Samuel G. Fadel
🏛️ University of Campinas | Leuphana University | Norwegian University of Science and Technology

Normalized flow generative models suffer from interpolation paths deviating from the data manifold, primarily due to norm drift induced by Gaussian base distributions in latent space. To address this, we propose a norm-constrained base distribution reconstruction framework—introducing Dirichlet and von Mises–Fisher distributions into normalized flows for the first time. These distributions explicitly constrain latent variables to the unit simplex or unit hypersphere, respectively, ensuring geometrically consistent interpolation trajectories. Our method requires no architectural modifications to the flow network and provides an interpretable, unambiguous interpolation criterion, effectively overcoming interpolation distortion inherent to the Gaussian assumption. Experiments demonstrate consistent improvements over baselines across all major evaluation metrics: bits/dim, Fréchet Inception Distance (FID), and Kernel Inception Distance (KID). Interpolation quality is significantly enhanced while strictly preserving original generation performance.

Addressing side effects of linear interpolation pathsEnabling principled interpolation through base distribution changesImproving interpolation in normalizing flow generative models

MUSO: Achieving Exact Machine Unlearning in Over-Parameterized Regimes

Oct 11, 2024
RY
Ruikai Yang
🏛️ Shanghai Jiao Tong University

Machine unlearning in over-parameterized models (e.g., neural networks) remains fundamentally limited—existing approaches achieve only approximate forgetting in output space, failing to guarantee exact parameter-space equivalence. Method: This paper establishes, for the first time, a theoretical proof that exact parameter-level unlearning is achievable in over-parameterized linear models via data relabeling. Building on this insight, we propose an alternating optimization framework that jointly optimizes relabeling and unlearning objectives, and extend it to nonlinear networks using random feature analysis, SGD dynamics modeling, and over-parameterization theory. Contribution/Results: Our method significantly outperforms state-of-the-art unlearning approaches—especially relabeling-based ones—across diverse benchmarks. Crucially, it provides the first empirical validation of exact parameter-space unlearning in practical neural networks, demonstrating both feasibility and effectiveness.

Achieving exact machine unlearning in over-parameterized modelsProposing an algorithm for unlearning and relabeling in nonlinear networksValidating relabeling and fine-tuning for parameter-space unlearning

Latest Papers

What's happening recently
View more

This work addresses the challenges posed by temporal, spatial, and conditional output distribution shifts in time series forecasting, which commonly undermine the effectiveness of normalization methods. For the first time, we systematically analyze the mechanism of reversible instance normalization (RevIN) and, through ablation studies, reveal that certain components are redundant or even detrimental to performance. Building on these insights, we propose a refined perspective that clearly distinguishes essential from non-essential elements of RevIN, leading to substantially improved model robustness and generalization. Our findings provide a principled foundation for designing more efficient and effective normalization strategies tailored specifically for time series data.

data normalizationdistribution shiftReversible Instance Normalization

Non-stationarity often degrades the predictive performance of causal large time series models, and conventional normalization methods may inadvertently introduce future information leakage during training. This work systematically evaluates multiple normalization strategies—including causal normalization and normalization based on statistics from initial observations—within a causal Transformer architecture, combined with time series chunking. It presents the first large-scale comparison of their impact on training stability and forecasting accuracy in autoregressive modeling. The study demonstrates that the choice of normalization critically determines model performance and offers key practical guidance for avoiding information leakage and enhancing the efficiency of causal time series modeling.

causal time-series modelsforecasting performanceinformation leakage

This work investigates the necessity of sample-dependent normalization in pre-normalized Transformers and proposes TaperNorm, a plug-and-play dynamic normalization alternative. TaperNorm initially mimics standard normalization during early training and then smoothly transitions to a sample-independent linear mapping via an EMA-calibrated global gating mechanism coupled with cosine annealing scheduling. The scaling parameters are fused into adjacent linear layers to enhance inference efficiency. This approach is the first to dynamically remove normalization during training without performance degradation, revealing that normalization’s primary role is to provide a scale anchor that prevents unbounded logit growth. A fixed-target auxiliary loss is introduced as a replacement. Experiments show that TaperNorm maintains training effectiveness while eliminating per-token statistics, achieving up to a 1.22× speedup in last-token inference throughput.

Efficient InferenceNormalizationSample-dependent Statistics

This work addresses key challenges in post-hoc calibration—namely nonlinear miscalibration, poor scalability to large numbers of classes, and perturbation of original predictions—by proposing Invertible Logit Transformation (InvLT). InvLT applies a shared-parameter scalar MLP element-wise to pre-softmax logits and incorporates a paired inverse network with soft monotonicity constraints. This design achieves high expressiveness and strong class scalability without introducing class-dependent parameters or requiring model retraining, while rigorously preserving the original classification accuracy. Extensive experiments across diverse image classification benchmarks and model architectures demonstrate that InvLT consistently outperforms existing calibration methods on standard calibration metrics, all while maintaining the original predictive performance without degradation.

classification accuracylogits transformationmonotonicity

This work investigates how to reconstruct input images from neural network outputs to uncover the features underlying model decisions. To this end, two novel inversion methods are proposed: a forward inversion approach leveraging the input Jacobian matrix combined with root-finding algorithms, and a backward inversion technique that iteratively inverts layer-by-layer while injecting random vectors into the nullspace of each layer’s linear transformation. For the first time, high-fidelity input reconstructions are achieved on Transformers and linear sequence networks. The generated images, though appearing random, consistently yield near-100% classification confidence and densely span the feasible input space. This approach substantially outperforms existing methods and effectively exposes the model’s reliance on non-semantic features and its inherent vulnerabilities.

Class-specific InputsInput ReconstructionInverse Problem

Hot Scholars

YD

Yun Dai

OpenAI
deep learningLLMML systemsdistributed training
CZ

Changqing Zhang

Professor, Tianjin University
Machine LearningMultimodal LearningLLM
RC

Rajen Chatterjee

Apple Inc
Machine TranslationAutomatic Post-EditingNatural Language ProcessingCrowdsourcing
DH

Dayananda Herurkar

Researcher at DFKI Germany
Outlier DetectionAnomaly DetectionNatural Language ProcessingFederated Learning