density score learning

Design and implement estimators that learn differentiable unnormalized density scores or density ratios by contrasting data with noise, using objectives such as noise-contrastive estimation or density-based scoring. Build and analyze discriminative training procedures (e.g., binary or k+1 discriminators) that produce score functions and density gradients usable for guiding generation, sampling, or likelihood-ratio decisions and evaluate their statistical and computational properties.

densityscorelearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.24
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This paper addresses the fragmented development and lack of theoretical unification across diverse methods for learning unnormalized distributions. To resolve this, it adopts Noise Contrastive Estimation (NCE) as a unifying statistical framework. The study systematically models and reveals the intrinsic consistency of several classical estimators—including Score Matching, Pseudo-likelihood, and others—within the NCE paradigm, constituting the first such unified characterization. Furthermore, for exponential-family unnormalized models under standard regularity conditions, the paper establishes the first tight finite-sample convergence rate theory, delivering an explicit upper bound on estimation error. Key theoretical contributions—including the NCE-based unification and the finite-sample analysis—are novel and advance the statistical foundations and interpretability of unnormalized models. The results bridge methodological gaps across domains and provide principled guidance for estimator design and analysis.

Connecting independently proposed methods through NCE frameworkFinite-sample convergence rates for exponential familiesUnified perspective on learning unnormalized distributions via NCE

Density Ratio Estimation-based Bayesian Optimization with Semi-Supervised Learning

May 24, 2023
JK
Jungtaek Kim
🏛️ University of Pittsburgh

In density-ratio estimation-based Bayesian optimization, supervised classifiers often exhibit overconfidence in known optimal candidates, leading to poor generalization and biased decision-making. Method: This paper proposes the first semi-supervised learning-enhanced framework for this setting, integrating density-ratio estimation with semi-supervised classification (e.g., Π-model or UDA) under a Bayesian sequential decision-making framework augmented with active sampling. It operates under low-label-budget conditions—requiring only a few labeled points while leveraging abundant unlabeled data (either randomly sampled or from a fixed pool). Contribution/Results: The key innovation lies in using unlabeled data to calibrate classifier confidence, effectively mitigating overfitting and discriminative bias. Experiments demonstrate that, under limited labeling budgets, the method achieves an average 23.7% acceleration in optimization speed and an 18.4% improvement in convergence accuracy over baseline approaches, validating its efficiency, robustness, and generalization capability.

Addressing overconfidence in supervised classifiers for optimizationEstimating density ratio for Bayesian optimization accuracyIncorporating semi-supervised learning with unlabeled data points

Rejection via Learning Density Ratios

May 29, 2024
AS
Alexander Soen
🏛️ Amazon

This work addresses the problem that classification models are forced to predict even under high uncertainty. We propose a rejection mechanism based on density ratio estimation (DRE), modeling rejection as estimating the ratio between the true data distribution and an idealized target distribution. Our method optimizes a risk function regularized by α-divergence—departing for the first time from conventional loss-augmentation paradigms. It endows rejection decisions with well-defined probabilistic semantics and distribution-level interpretability, and naturally accommodates pre-trained models. Empirically, it significantly improves both rejection accuracy and classification reliability on both clean and noisy datasets. Key contributions are: (1) the first formalization of rejection learning as a DRE problem; (2) the introduction of φ-divergence regularization—specifically the α-divergence family—to achieve distributionally robust optimization; and (3) a theoretically rigorous yet practical framework for interpretable rejection.

Develops a method for classification with rejection using density ratiosProposes an idealized data distribution to enhance model performanceUtilizes $alpha$-divergence for regularization and rejection decisions

Nonlinear denoising score matching for enhanced learning of structured distributions

May 24, 2024
JB
Jeremiah Birrell
🏛️ Texas State University | University of Massachusetts Amherst

To address the limitations of score-based generative models in modeling structured distributions—such as multimodal or approximately symmetric ones—and their poor generalization under small-sample regimes, this paper proposes the Nonlinear Denoising Score Matching (NDSM) framework. NDSM introduces learnable nonlinear drift into score matching for the first time, leveraging nonlinear stochastic differential equations (SDEs) to enhance structural representation capability. It further incorporates a neural control variate technique to substantially reduce gradient estimation variance. Crucially, NDSM enables data-driven structural embedding without requiring explicit symmetry priors. Experiments demonstrate that NDSM effectively mitigates mode collapse on both low-dimensional structured distributions and high-dimensional image data, improves small-sample generalization, and accurately captures approximate symmetries—outperforming equivariant networks and linear score matching baselines.

Addresses training challenges with nonlinear denoising score matchingEnhances performance in high-dimensional data with reduced computational costImproves learning of structured distributions using nonlinear noising dynamics

Alpha-divergence loss function for neural density ratio estimation

Feb 03, 2024
YK
Yoshiaki Kitazawa
🏛️ NTT DATA Mathematical Systems Inc.

Existing neural density ratio estimation (DRE) methods employing KL-divergence loss suffer from overfitting and training instability due to the loss’s unboundedness below, vanishing gradients, mini-batch bias, and sensitivity to sample size. This work introduces, for the first time, the α-divergence family into DRE loss design. Leveraging its f-divergence variational representation, we construct the α-Div loss—a bounded, differentiable objective that simultaneously ensures loss boundedness and gradient stability. The density ratio is parameterized by a neural network and optimized via stochastic gradient descent. Experiments demonstrate that α-Div significantly improves training stability and convergence speed, maintaining effective optimization even under high KL divergence between distributions. Its root-mean-square error (RMSE) accuracy matches that of KL-loss-based methods, indicating that the fundamental accuracy limit stems from intrinsic data properties rather than the choice of loss function.

Addresses optimization challenges in density ratio estimation.Explores impact of α-divergence on DRE accuracy.Proposes α-divergence loss function for stable DRE optimization.

Latest Papers

What's happening recently
View more

Diffusion models are prone to reproducing training samples during generation due to memorization, which compromises their generalization capability. This work theoretically shows that the empirical score function consists of Gaussian scores weighted by a sharp softmax, causing individual training samples to dominate the generation process. To address this, the authors propose two novel techniques: Noise Unconditioning and Temperature Smoothing. The former adaptively adjusts sample weights, while the latter explicitly controls the softmax temperature; together, they yield a smoothed approximation of the score function, ensuring that sampling is guided by the local data manifold rather than isolated points. Experimental results validate the theoretical analysis, demonstrating that the proposed methods significantly improve generalization across multiple datasets while preserving high-quality generation.

diffusion modelsgeneralizationmemorization

This work addresses the challenge of performing Bayesian inference in unnormalized models, where the intractable normalization constant hinders conventional approaches. The authors propose a fully Bayesian inference framework that treats the normalization constant as an unknown parameter and reformulates the inference problem as a binary classification task between observed data and noise samples, leveraging noise contrastive estimation. By integrating Pólya–Gamma data augmentation with Gibbs sampling, the method efficiently handles exponential-family unnormalized models without requiring tuning parameters, thereby circumventing the sensitivity to likelihood tempering that plagues existing techniques. Experiments on time-varying density point processes and sparse toroidal graphical models demonstrate that the approach yields accurate parameter estimates and reliable uncertainty quantification.

Bayesian inferenceenergy-based modelsintractable likelihood

该研究通过密度比重评分(DRR)方法解决不平衡分类问题,利用调查整合法对多数样本重新加权,并结合基础分类器提高稀有类排名的精度。

Density-Ratio RescoringDual ScoreImbalanced Classification

This work addresses the challenging problem of covariate shift adaptation when the density ratio is unbounded—a setting where existing methods often rely on unrealistic assumptions that the density ratio is either bounded or exactly known. To overcome this limitation, the authors propose a novel three-step estimation procedure: first estimating a relative density ratio, then applying truncation to control its unboundedness, and finally transforming it into a standard density ratio to serve as importance weights in regression. This approach is the first to directly tackle unbounded density ratios, establishing non-asymptotic convergence guarantees that achieve minimax-optimal or near-optimal rates for both the density ratio and the regression function. The method significantly enhances both the theoretical rigor and empirical performance of covariate shift adaptation under realistic conditions.

covariate shift adaptationdensity ratio estimationimportance weighting

Learning density ratios in causal inference using Bregman-Riesz regression

Oct 17, 2025
OJ
Oliver J. Hines
🏛️ Columbia University

This paper addresses the instability and curse of dimensionality in density-ratio estimation for high-dimensional causal inference. We propose a direct density-ratio learning method grounded in a unified framework integrating Bregman divergences and the Riesz representation theorem. Unlike conventional two-stage density estimation, our approach formulates the density ratio as a regression problem under the Riesz representation, using a Bregman divergence as the loss function, and incorporates classification-inspired objectives and data augmentation to mitigate estimation bias under unobserved intervention distributions. To our knowledge, this is the first work to theoretically unify Bregman divergences, Riesz regression, and probabilistic classification, while supporting diverse model classes—including gradient boosting, neural networks, and kernel methods. Extensive simulations demonstrate that different Bregman divergences and augmentation strategies enhance robustness. A publicly available Python package enables flexible implementation.

Addressing instability in density estimation with high-dimensional covariatesDeveloping practical tools for density ratio learning using various ML modelsUnifying density ratio estimation methods for causal inference applications

Hot Scholars

SE

Stefano Ermon

Stanford University
Artificial IntelligenceMachine Learning
MX

Minkai Xu

Stanford University
Generative AI
YL

Yann LeCun

Chief AI Scientist at Facebook & JT Schwarz Professor at the Courant Institute, New York University
AImachine learningcomputer visionrobotics
WT

Wei Tsang Ooi

National University of Singapore
Multimedia SystemsInteractive SystemsIntelligent Systems