density estimation

Design and implement methods to estimate probability density functions or local density measures from data, using parametric, nonparametric, neural, or spectral approaches and unsupervised learning; this includes constructing and training neural density estimators and nonparametric estimators. Build and analyze iterative or adaptive ‘densification’ and sampling policies that allocate a fixed computational or sample budget to refine local estimates (for example by inserting local components where sampling error concentrates), and validate estimators with statistical goodness-of-fit tests and measures such as local false discovery rates.

densityestimation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.48
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$205K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Existing neural density estimation methods exhibit strong empirical performance but struggle to simultaneously enforce non-negativity and unit-mass constraints, and lack theoretical guarantees for adaptive convergence on low-dimensional structured densities. This paper proposes a structure-agnostic, classification-induced neural density estimation framework: by reformulating density ratio estimation as a binary classification task, it implicitly satisfies probabilistic constraints without explicit regularization. We establish theoretical guarantees showing that, under high-dimensional densities supported on low-dimensional manifolds, the estimator achieves adaptive convergence rates strictly faster than standard nonparametric rates. Moreover, the framework naturally yields an efficient sampling mechanism. Extensive experiments on synthetic and real-world datasets demonstrate significant improvements in convergence speed, density estimation accuracy, and sample quality. To our knowledge, this work provides the first rigorous theoretical connection between the classification paradigm and neural density estimation.

Achieves adaptive convergence rates for low-dimensional structured densitiesIntegrates with generative sampling pipelines like diffusion models for faster convergenceProposes a neural density estimator with easy implementation and theoretical guarantees

Transforming Conditional Density Estimation Into a Single Nonparametric Regression Task

Nov 23, 2025
AG
Alexander G. Reisach
🏛️ CNRS | Université Paris Cité | MODAL'X | Université Paris Nanterre | Harvard University

This paper addresses the challenge of applying powerful regression models directly to high-dimensional conditional density estimation (CDE). We propose a novel framework that reformulates CDE as a single nonparametric regression task by constructing labeled auxiliary samples, thereby mapping density estimation into a standard regression problem. Our approach imposes no parametric assumptions on the conditional density and enables plug-and-play integration of arbitrary state-of-the-art regressors—including deep neural networks and gradient-boosted trees. We establish theoretical consistency: the proposed estimator converges almost surely to the true conditional density as the sample size tends to infinity. Extensive experiments on synthetic data, U.S. Census records, and satellite imagery demonstrate that our method significantly outperforms existing CDE benchmarks across diverse real-world scenarios. Moreover, the results align with domain-specific prior knowledge, underscoring the method’s balance of theoretical rigor and practical efficacy.

Achieves state-of-the-art performance on survey and satellite datasetsLeverages neural networks for high-dimensional density estimationTransforms conditional density estimation into nonparametric regression

This study addresses the challenge of enhancing density estimation performance by integrating the structural advantages of parametric models while preserving the flexibility of nonparametric methods. The authors propose a locally parametrized nonparametric density estimator that, for each point \(x\), estimates an optimal local parameter \(\hat{\theta}(x)\) via kernel-smoothed likelihood, yielding an estimator of the form \(f(x, \hat{\theta}(x))\). This approach achieves near-full-likelihood efficiency under correct model specification and retains nonparametric robustness under misspecification, effectively serving as a semiparametric realization of higher-order kernel methods. Theoretical analysis and empirical experiments demonstrate that the proposed estimator exhibits variance comparable to classical kernel density estimation but substantially reduced bias, leading to significantly improved accuracy in neighborhoods of well-specified parametric models.

density estimationlocal likelihoodnonparametric

Copula Density Neural Estimation

Nov 25, 2022
NA
N. A. Letizia
🏛️ University of Klagenfurt

This work addresses probabilistic density estimation for high-dimensional complex data. We propose a neural Copula modeling framework that explicitly decouples marginal distributions from dependency structures—a departure from conventional joint modeling approaches. Integrating Copula theory with deep neural networks, our method employs a normalizing flow architecture to separately parameterize marginals and the high-dimensional copula function, enabling differentiable density evaluation, generative sampling, and tractable inference. Leveraging marginal-joint decoupled training and a differentiable mutual information estimator, the model achieves superior performance over kernel density estimation and state-of-the-art neural density estimators across diverse complex distributions. Empirical evaluation demonstrates its effectiveness in high-accuracy mutual information estimation and photorealistic data generation, validating both theoretical design and practical utility.

Estimating probability density from observed dataModeling long-range dependencies in big dataSeparating univariate marginals from joint dependence structure

A survey of Monte Carlo methods for noisy and costly densities with application to reinforcement learning

Aug 01, 2021
FL
F. Llorente
🏛️ Stony Brook University | Universitá degli Studi di Catania | École Polytechnique | Universidad Carlos III de Madrid

Addressing the challenges of expensive, stochastic, and analytically intractable function evaluations in reinforcement learning (RL) and approximate Bayesian computation (ABC), this paper systematically reviews and refactors the Monte Carlo methodology framework. We first unify surrogate modeling approaches—designed for costly, noisy, and intractable densities—into three principled categories, and propose a modular surrogate modeling paradigm that jointly optimizes accuracy, computational cost, and robustness. Our framework is innovatively extended to likelihood-free inference and online RL settings. Integrating Bayesian optimization, Gaussian processes, sequential Monte Carlo, importance sampling, and adaptive experimental design, we conduct comprehensive numerical experiments to quantitatively characterize the trade-offs among sample efficiency, convergence stability, and noise robustness. The results provide a reusable, principled guideline for method selection in RL policy evaluation and hyperparameter optimization.

Approximate Bayesian ComputationMonte Carlo MethodsReinforcement Learning

Latest Papers

What's happening recently
View more

This work proposes a semiparametric density estimation approach that addresses the inefficiency of traditional kernel density estimation in high-dimensional settings or when the underlying distribution deviates substantially from assumed parametric forms. The method multiplies a parametric initial guess—such as a normal distribution—by a nonparametric kernel-based correction factor, thereby preserving robustness against departures from the parametric family while significantly enhancing local estimation accuracy. By integrating parametric priors with nonparametric adjustments, the framework incorporates a tailored bandwidth selection strategy and naturally extends to nonparametric regression. Theoretical analysis and extensive simulations demonstrate that, even when the true density markedly departs from normality, the proposed estimator consistently outperforms conventional kernel density estimators across various Gaussian mixture models, particularly excelling in high-dimensional scenarios.

density estimationkernel estimatornonparametric

This study addresses the well-known boundary bias problem in existing nonparametric methods for estimating quantile density functions and their derivatives. To overcome this limitation, the authors propose a novel local polynomial-based estimation strategy that significantly enhances performance near boundaries. The proposed estimator demonstrates superior properties in terms of bias reduction, asymptotic variance, and boundary behavior compared to classical approaches. Moreover, the paper establishes the asymptotic normality of the estimator through rigorous theoretical analysis, thereby providing a more reliable foundation for nonparametric inference on both quantile densities and their derivatives. This advancement offers improved theoretical guarantees and practical utility for applications requiring accurate estimation in boundary regions.

boundary propertieslocal polynomial estimationnonparametric estimation

This study addresses the challenges of non-negativity, normalization, and accuracy in estimating probability densities from empirical characteristic functions over a fixed time window. The authors propose a neural network approach trained directly in the Fourier domain, leveraging the closed-form characteristic function of a Gaussian–Laplace mixture model to enforce non-negativity and unit integral of the resulting density. The method is applicable to both independent and identically distributed data and dependent data resampling scenarios. As the first work to integrate neural networks with Fourier-domain training for density estimation, it derives a multidimensional L₂ error bound accounting for truncation, empirical, discretization, and sampling errors. Experiments demonstrate that the method matches the EM algorithm on Gaussian mixture benchmarks, substantially outperforms existing approaches on heavy-tailed distributions, exhibits L₂ error decay consistent with theoretical predictions, and successfully estimates the annual return distribution of the Australian stock market.

characteristic functiondensity estimationdependent data

This study addresses the fundamental statistical problem of conditional density estimation by systematically comparing classical nonparametric approaches—such as single-index models, basis expansion methods (e.g., FlexCode), and DeepCDE—with modern generative models, including conditional GANs and conditional denoising diffusion probabilistic models. For the first time, these methods are evaluated within a unified and reproducible framework using metrics like mean squared error and Wasserstein distance to assess their accuracy, flexibility, and computational cost in estimating conditional means and standard deviations. The analysis clarifies the performance boundaries and practical applicability of each approach, offering both theoretical insights and actionable guidance for selecting appropriate methods in predictive modeling, uncertainty quantification, and probabilistic inference tasks.

conditional distribution estimationgenerative modelsnonparametric methods

This work addresses the computational expense and poor sample efficiency associated with modeling high-dimensional, non-Gaussian probability density functions in nonlinear dynamical systems. To overcome these challenges, the authors propose an efficient estimation approach based on a semi-nonparametric (SNP) density model. The method constructs a strictly positive density representation using Hermite polynomial basis functions and integrates Monte Carlo integration with a convex relaxation optimization strategy to significantly enhance the accuracy and stability of density and quantile estimation under limited sample sizes. Experimental results on the Lorenz chaotic system demonstrate that the proposed method accurately captures complex non-Gaussian structures and reliably computes quantiles using substantially fewer samples than conventional Monte Carlo techniques.

chaotic systemsdata efficiencynon-Gaussian density estimation

Hot Scholars

SX

Shuyin Xia

Professor, School of Computer Science, Chongqing University of Posts and Telecommunications
Granular computingClusteringRough setsClassifiers
SL

Sergey Levine

UC Berkeley, Physical Intelligence
Machine LearningRoboticsReinforcement Learning
JM

José Miguel Hernández-Lobato

Professor of Machine Learning, University of Cambridge
Bayesian deep learningapproximate inferencedeep generative modelingautomatic molecular design
SK

Samuel Kaski

Director, ELLIS Institute Finland; Professor, Aalto University and University of Manchester
Probabilistic machine learningAI4ScienceCollaborative AI
ML

Morteza Lahijanian

University of Colorado Boulder
Safe AIformal methodsstochastic systemsmotion planning