Score
Design and implement estimators that learn differentiable unnormalized density scores or density ratios by contrasting data with noise, using objectives such as noise-contrastive estimation or density-based scoring. Build and analyze discriminative training procedures (e.g., binary or k+1 discriminators) that produce score functions and density gradients usable for guiding generation, sampling, or likelihood-ratio decisions and evaluate their statistical and computational properties.
This paper addresses the fragmented development and lack of theoretical unification across diverse methods for learning unnormalized distributions. To resolve this, it adopts Noise Contrastive Estimation (NCE) as a unifying statistical framework. The study systematically models and reveals the intrinsic consistency of several classical estimators—including Score Matching, Pseudo-likelihood, and others—within the NCE paradigm, constituting the first such unified characterization. Furthermore, for exponential-family unnormalized models under standard regularity conditions, the paper establishes the first tight finite-sample convergence rate theory, delivering an explicit upper bound on estimation error. Key theoretical contributions—including the NCE-based unification and the finite-sample analysis—are novel and advance the statistical foundations and interpretability of unnormalized models. The results bridge methodological gaps across domains and provide principled guidance for estimator design and analysis.
In density-ratio estimation-based Bayesian optimization, supervised classifiers often exhibit overconfidence in known optimal candidates, leading to poor generalization and biased decision-making. Method: This paper proposes the first semi-supervised learning-enhanced framework for this setting, integrating density-ratio estimation with semi-supervised classification (e.g., Π-model or UDA) under a Bayesian sequential decision-making framework augmented with active sampling. It operates under low-label-budget conditions—requiring only a few labeled points while leveraging abundant unlabeled data (either randomly sampled or from a fixed pool). Contribution/Results: The key innovation lies in using unlabeled data to calibrate classifier confidence, effectively mitigating overfitting and discriminative bias. Experiments demonstrate that, under limited labeling budgets, the method achieves an average 23.7% acceleration in optimization speed and an 18.4% improvement in convergence accuracy over baseline approaches, validating its efficiency, robustness, and generalization capability.
This work addresses the problem that classification models are forced to predict even under high uncertainty. We propose a rejection mechanism based on density ratio estimation (DRE), modeling rejection as estimating the ratio between the true data distribution and an idealized target distribution. Our method optimizes a risk function regularized by α-divergence—departing for the first time from conventional loss-augmentation paradigms. It endows rejection decisions with well-defined probabilistic semantics and distribution-level interpretability, and naturally accommodates pre-trained models. Empirically, it significantly improves both rejection accuracy and classification reliability on both clean and noisy datasets. Key contributions are: (1) the first formalization of rejection learning as a DRE problem; (2) the introduction of φ-divergence regularization—specifically the α-divergence family—to achieve distributionally robust optimization; and (3) a theoretically rigorous yet practical framework for interpretable rejection.
To address the limitations of score-based generative models in modeling structured distributions—such as multimodal or approximately symmetric ones—and their poor generalization under small-sample regimes, this paper proposes the Nonlinear Denoising Score Matching (NDSM) framework. NDSM introduces learnable nonlinear drift into score matching for the first time, leveraging nonlinear stochastic differential equations (SDEs) to enhance structural representation capability. It further incorporates a neural control variate technique to substantially reduce gradient estimation variance. Crucially, NDSM enables data-driven structural embedding without requiring explicit symmetry priors. Experiments demonstrate that NDSM effectively mitigates mode collapse on both low-dimensional structured distributions and high-dimensional image data, improves small-sample generalization, and accurately captures approximate symmetries—outperforming equivariant networks and linear score matching baselines.
Existing neural density ratio estimation (DRE) methods employing KL-divergence loss suffer from overfitting and training instability due to the loss’s unboundedness below, vanishing gradients, mini-batch bias, and sensitivity to sample size. This work introduces, for the first time, the α-divergence family into DRE loss design. Leveraging its f-divergence variational representation, we construct the α-Div loss—a bounded, differentiable objective that simultaneously ensures loss boundedness and gradient stability. The density ratio is parameterized by a neural network and optimized via stochastic gradient descent. Experiments demonstrate that α-Div significantly improves training stability and convergence speed, maintaining effective optimization even under high KL divergence between distributions. Its root-mean-square error (RMSE) accuracy matches that of KL-loss-based methods, indicating that the fundamental accuracy limit stems from intrinsic data properties rather than the choice of loss function.
Diffusion models are prone to reproducing training samples during generation due to memorization, which compromises their generalization capability. This work theoretically shows that the empirical score function consists of Gaussian scores weighted by a sharp softmax, causing individual training samples to dominate the generation process. To address this, the authors propose two novel techniques: Noise Unconditioning and Temperature Smoothing. The former adaptively adjusts sample weights, while the latter explicitly controls the softmax temperature; together, they yield a smoothed approximation of the score function, ensuring that sampling is guided by the local data manifold rather than isolated points. Experimental results validate the theoretical analysis, demonstrating that the proposed methods significantly improve generalization across multiple datasets while preserving high-quality generation.
This work addresses the challenge of performing Bayesian inference in unnormalized models, where the intractable normalization constant hinders conventional approaches. The authors propose a fully Bayesian inference framework that treats the normalization constant as an unknown parameter and reformulates the inference problem as a binary classification task between observed data and noise samples, leveraging noise contrastive estimation. By integrating Pólya–Gamma data augmentation with Gibbs sampling, the method efficiently handles exponential-family unnormalized models without requiring tuning parameters, thereby circumventing the sensitivity to likelihood tempering that plagues existing techniques. Experiments on time-varying density point processes and sparse toroidal graphical models demonstrate that the approach yields accurate parameter estimates and reliable uncertainty quantification.
该研究通过密度比重评分(DRR)方法解决不平衡分类问题,利用调查整合法对多数样本重新加权,并结合基础分类器提高稀有类排名的精度。
This work addresses the challenging problem of covariate shift adaptation when the density ratio is unbounded—a setting where existing methods often rely on unrealistic assumptions that the density ratio is either bounded or exactly known. To overcome this limitation, the authors propose a novel three-step estimation procedure: first estimating a relative density ratio, then applying truncation to control its unboundedness, and finally transforming it into a standard density ratio to serve as importance weights in regression. This approach is the first to directly tackle unbounded density ratios, establishing non-asymptotic convergence guarantees that achieve minimax-optimal or near-optimal rates for both the density ratio and the regression function. The method significantly enhances both the theoretical rigor and empirical performance of covariate shift adaptation under realistic conditions.
This paper addresses the instability and curse of dimensionality in density-ratio estimation for high-dimensional causal inference. We propose a direct density-ratio learning method grounded in a unified framework integrating Bregman divergences and the Riesz representation theorem. Unlike conventional two-stage density estimation, our approach formulates the density ratio as a regression problem under the Riesz representation, using a Bregman divergence as the loss function, and incorporates classification-inspired objectives and data augmentation to mitigate estimation bias under unobserved intervention distributions. To our knowledge, this is the first work to theoretically unify Bregman divergences, Riesz regression, and probabilistic classification, while supporting diverse model classes—including gradient boosting, neural networks, and kernel methods. Extensive simulations demonstrate that different Bregman divergences and augmentation strategies enhance robustness. A publicly available Python package enables flexible implementation.