density-ratio estimation

Designs, implements, and analyzes estimators that directly recover the ratio between two probability densities or between class-conditional densities from samples — including classifier-based likelihood-ratio methods, nonparametric kernel and empirical-likelihood approaches, spectral/log-density-ratio estimators built from covariance eigenspectra, and importance-weight estimation. Work covers building offline estimators, choosing regularization and kernel/hyperparameters (e.g., ridge or least-squares across spectral components, kernel bandwidth and gating), quantifying bias/variance in low-sample regimes, and producing ratios for tasks such as importance weighting, thresholded classification, or likelihood-based decision rules.

density-ratioestimation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.16
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Likelihood Ratio Tests by Kernel Gaussian Embedding

Aug 11, 2025
LV
Leonardo V. Santoro
🏛️ École Polytechnique Fédérale de Lausanne

This paper addresses the nonparametric two-sample testing problem by proposing a likelihood ratio test based on kernel Gaussian embeddings. The method jointly embeds the two distributions into a reproducing kernel Hilbert space (RKHS) via their kernel mean and kernel covariance operators, mapping them to mutually singular Gaussian measures. A regularized likelihood ratio statistic—constructed using the relative entropy between these Gaussian embeddings—is shown to converge to zero under the null hypothesis and diverge to infinity under the alternative, enabling natural hypothesis separation. The approach integrates permutation-based calibration and spectral regularization to ensure finite-sample stability. Theoretically, the test is proven to be consistent and possesses uniform power bounds. Empirically, it significantly outperforms state-of-the-art methods—including MMD and C2ST—in high-dimensional, weak-signal regimes, unifying and extending the kernel embedding-based testing framework.

Addresses high-dimensional weak-signal regime testingDetects distribution equality versus singularityDevelops kernel-based nonparametric two-sample test

Estimating Unbounded Density Ratios: Applications in Error Control under Covariate Shift

Mar 29, 2025
SX
Shuntuo Xu
🏛️ East China Normal University | The Hong Kong Polytechnic University

This paper addresses density ratio estimation under covariate shift without assuming boundedness—a departure from conventional bounded-density-ratio assumptions. Methodologically, it establishes novel upper bounds on estimation error for unbounded domains and ranges, leveraging least-squares and logistic regression losses, nonparametric estimation, and minimax analysis to derive optimal convergence rates up to logarithmic factors. Theoretically, it is the first to reveal that tail decay behavior of the density ratio fundamentally governs cross-domain generalization error, and identifies a sufficient condition—requiring no loss correction—for controlled generalization error. Moreover, it proves that, on suitably chosen target domains, source-domain estimators can outperform correction-based methods relying on the true density ratio. Extensive simulations validate the theoretical findings. The results provide a rigorous, error-controlled generalization theory for nonparametric regression and conditional flow models under covariate shift.

Analyzing tail properties under covariate shift conditionsComparing source and loss-corrected estimator performanceEstimating unbounded density ratios for error control

Demystifying Spectral Bias on Real-World Data

Jun 04, 2024
IL
Itay Lavie
🏛️ Hebrew University of Jerusalem

Kernel ridge regression (KRR) and Gaussian processes (GPs) suffer from degraded learnability on real-world data due to rapid eigenvalue decay of the kernel matrix, impairing their ability to recover target functions. Method: We propose a spectral learnability analysis framework that introduces the idealized data measure’s eigen-spectrum as a theoretical benchmark, quantifying spectral deviation of real data from this benchmark. Leveraging symmetry properties of the kernel operator, we derive tight theoretical bounds on learnability, revealing bottlenecks in learning components aligned with high-eigenvalue eigenvectors. The framework integrates kernel spectral analysis, eigenvalue estimation, and symmetry-driven generalization theory to quantitatively characterize the learnability of kernel methods under realistic data distributions. Results: Experiments demonstrate substantial improvements in predictive accuracy for learnability estimation. The framework establishes a novel paradigm for theoretically interpreting and practically deploying kernel methods on non-ideal, real-world data.

Cross-dataset learnability analysisEigenvalue challenges in real-world dataSpectral bias in complex datasets

Spectral Estimators for Structured Generalized Linear Models via Approximate Message Passing

Aug 28, 2023
YZ
Yihan Zhang
🏛️ Institute of Science and Technology Austria | University of Cambridge

Parameter estimation in high-dimensional structured generalized linear models suffers from low efficiency, particularly under realistic design matrices exhibiting anisotropy and strong correlations. Method: This paper introduces a novel spectral estimation framework based on Approximate Message Passing (AMP). Contribution/Results: We provide the first exact asymptotic characterization of spectral estimators under correlated Gaussian designs. We identify a universally optimal covariance-adaptive preprocessing strategy, partially resolving a long-standing conjecture on optimal spectral estimation for rotationally invariant models. Theoretically and empirically, our approach substantially reduces sample complexity and achieves provably statistically optimal estimation accuracy—outperforming existing heuristic methods on canonical designs from computational imaging and genomics.

Characterizing spectral estimators for correlated Gaussian designsEstimating parameters in high-dimensional generalized linear modelsIdentifying optimal preprocessing for efficient parameter estimation

Latest Papers

What's happening recently
View more

This work addresses the absence of closed-form solutions for relative log-densities in linearly parameterized probabilistic models—including unnormalized and conditional models—by proposing a closed-form spectral estimator that relies solely on first- and second-order feature moments. By expressing the Kullback–Leibler divergence as an integral of weighted chi-squared divergences, the problem is recast as a family of least-squares problems, yielding explicit spectral formulas applicable to a broad class of f-divergences. These formulas enable direct estimation of both divergences and log-density potentials. The approach naturally integrates with kernel methods or neural networks for feature learning, enjoys theoretical consistency guarantees, and demonstrates substantial empirical advantages over optimization-based variational methods—such as logistic and softmax regression—on synthetic data.

f-divergencesKL divergencelinearly parameterized probabilistic models

This study addresses the frequent misinterpretation of fluctuations in dominant eigensubspaces and scalar spectral functionals—such as the absorption ratio—in rolling covariance estimation as genuine market structural changes, when they are often artifacts of estimation noise, particularly under shrinkage. By leveraging perturbation analysis and calibrated inference, the work derives, for the first time, the first-order null distribution of eigensubspace variation under overlapping windows and establishes its invariance under rotation-equivariant shrinkage estimators. It further shows that only scale-invariant spectral functionals enjoy first-order immunity to elliptical kurtosis. To correct high-dimensional bias in the absorption ratio, a trace-preserving spiked debiased estimator is proposed. Theoretical results, supported by Davis–Kahan bounds, distribution-free confidence bands, and an estimator-aware bootstrap, are validated through simulations and successfully applied to equity data for reliable detection of true market structural shifts.

calibrated inferenceeigenspace perturbationerror propagation

This work addresses the challenges of bandwidth selection in kernel density estimation and its poor performance on multimodal, spiky, or discrete data by proposing a spectral method operating in the characteristic function domain. Treating bandwidth as a spectral truncation point, the approach integrates a data-driven noise-floor estimator, an adaptive Wiener tapering kernel, and a reconstruction architecture that combines a smooth basis with a band-limited residual. This framework unifies spectral decomposition, Wiener filtering, and Gaussian mixture modeling to enable automatic bandwidth selection and high-fidelity density estimation. Evaluated on the Marron–Wand benchmark suite (n = 5000), the method achieves the lowest integrated mean squared error and demonstrates superior robustness and reproducibility across six diverse real-world and synthetic datasets.

bandwidth selectionbias-variance tradeoffcharacteristic function

This study addresses the problem of regularized log-density ratio estimation under a high-dimensional Gaussian location model with ridge regularization. Focusing on the limited-sample regime, the authors construct two estimators—one based on variational methods and the other on spectral techniques—and, for the first time, derive their deterministic equivalents within a high-dimensional asymptotic framework. By integrating the convex Gaussian min-max theorem, resolvent analysis from random matrix theory, and variational characterizations of KL divergence, the theoretical analysis reveals that the variational estimator achieves lower risk in large-sample settings, whereas the spectral estimator is superior in small-sample scenarios due to its reduced variance. Numerical experiments validate the critical roles of optimal regularization strategies and the signal-to-dimension ratio in estimator performance, and further demonstrate that incorporating nuclear norm penalization enhances feature learning capabilities.

feature learningGaussian location modelhigh-dimensional asymptotics

This work addresses the challenging problem of covariate shift adaptation when the density ratio is unbounded—a setting where existing methods often rely on unrealistic assumptions that the density ratio is either bounded or exactly known. To overcome this limitation, the authors propose a novel three-step estimation procedure: first estimating a relative density ratio, then applying truncation to control its unboundedness, and finally transforming it into a standard density ratio to serve as importance weights in regression. This approach is the first to directly tackle unbounded density ratios, establishing non-asymptotic convergence guarantees that achieve minimax-optimal or near-optimal rates for both the density ratio and the regression function. The method significantly enhances both the theoretical rigor and empirical performance of covariate shift adaptation under realistic conditions.

covariate shift adaptationdensity ratio estimationimportance weighting

Hot Scholars

MK

Masahiro Kato

Mizuho-DL Financial Technology Co., Ltd. / The University of Tokyo
Economics
PL

Pengfei Li

University of Waterloo
Mixture model and biased sampling
SC

Seungjin Choi

Research Director of Machine Learning Lab, Intellicode
machine learning
SL

Sergey Levine

UC Berkeley, Physical Intelligence
Machine LearningRoboticsReinforcement Learning