Score
Design and implement methods to estimate probability density functions or local density measures from data, using parametric, nonparametric, neural, or spectral approaches and unsupervised learning; this includes constructing and training neural density estimators and nonparametric estimators. Build and analyze iterative or adaptive ‘densification’ and sampling policies that allocate a fixed computational or sample budget to refine local estimates (for example by inserting local components where sampling error concentrates), and validate estimators with statistical goodness-of-fit tests and measures such as local false discovery rates.
Existing neural density estimation methods exhibit strong empirical performance but struggle to simultaneously enforce non-negativity and unit-mass constraints, and lack theoretical guarantees for adaptive convergence on low-dimensional structured densities. This paper proposes a structure-agnostic, classification-induced neural density estimation framework: by reformulating density ratio estimation as a binary classification task, it implicitly satisfies probabilistic constraints without explicit regularization. We establish theoretical guarantees showing that, under high-dimensional densities supported on low-dimensional manifolds, the estimator achieves adaptive convergence rates strictly faster than standard nonparametric rates. Moreover, the framework naturally yields an efficient sampling mechanism. Extensive experiments on synthetic and real-world datasets demonstrate significant improvements in convergence speed, density estimation accuracy, and sample quality. To our knowledge, this work provides the first rigorous theoretical connection between the classification paradigm and neural density estimation.
This paper addresses the challenge of applying powerful regression models directly to high-dimensional conditional density estimation (CDE). We propose a novel framework that reformulates CDE as a single nonparametric regression task by constructing labeled auxiliary samples, thereby mapping density estimation into a standard regression problem. Our approach imposes no parametric assumptions on the conditional density and enables plug-and-play integration of arbitrary state-of-the-art regressors—including deep neural networks and gradient-boosted trees. We establish theoretical consistency: the proposed estimator converges almost surely to the true conditional density as the sample size tends to infinity. Extensive experiments on synthetic data, U.S. Census records, and satellite imagery demonstrate that our method significantly outperforms existing CDE benchmarks across diverse real-world scenarios. Moreover, the results align with domain-specific prior knowledge, underscoring the method’s balance of theoretical rigor and practical efficacy.
This study addresses the challenge of enhancing density estimation performance by integrating the structural advantages of parametric models while preserving the flexibility of nonparametric methods. The authors propose a locally parametrized nonparametric density estimator that, for each point \(x\), estimates an optimal local parameter \(\hat{\theta}(x)\) via kernel-smoothed likelihood, yielding an estimator of the form \(f(x, \hat{\theta}(x))\). This approach achieves near-full-likelihood efficiency under correct model specification and retains nonparametric robustness under misspecification, effectively serving as a semiparametric realization of higher-order kernel methods. Theoretical analysis and empirical experiments demonstrate that the proposed estimator exhibits variance comparable to classical kernel density estimation but substantially reduced bias, leading to significantly improved accuracy in neighborhoods of well-specified parametric models.
This work addresses probabilistic density estimation for high-dimensional complex data. We propose a neural Copula modeling framework that explicitly decouples marginal distributions from dependency structures—a departure from conventional joint modeling approaches. Integrating Copula theory with deep neural networks, our method employs a normalizing flow architecture to separately parameterize marginals and the high-dimensional copula function, enabling differentiable density evaluation, generative sampling, and tractable inference. Leveraging marginal-joint decoupled training and a differentiable mutual information estimator, the model achieves superior performance over kernel density estimation and state-of-the-art neural density estimators across diverse complex distributions. Empirical evaluation demonstrates its effectiveness in high-accuracy mutual information estimation and photorealistic data generation, validating both theoretical design and practical utility.
Addressing the challenges of expensive, stochastic, and analytically intractable function evaluations in reinforcement learning (RL) and approximate Bayesian computation (ABC), this paper systematically reviews and refactors the Monte Carlo methodology framework. We first unify surrogate modeling approaches—designed for costly, noisy, and intractable densities—into three principled categories, and propose a modular surrogate modeling paradigm that jointly optimizes accuracy, computational cost, and robustness. Our framework is innovatively extended to likelihood-free inference and online RL settings. Integrating Bayesian optimization, Gaussian processes, sequential Monte Carlo, importance sampling, and adaptive experimental design, we conduct comprehensive numerical experiments to quantitatively characterize the trade-offs among sample efficiency, convergence stability, and noise robustness. The results provide a reusable, principled guideline for method selection in RL policy evaluation and hyperparameter optimization.
This work proposes a semiparametric density estimation approach that addresses the inefficiency of traditional kernel density estimation in high-dimensional settings or when the underlying distribution deviates substantially from assumed parametric forms. The method multiplies a parametric initial guess—such as a normal distribution—by a nonparametric kernel-based correction factor, thereby preserving robustness against departures from the parametric family while significantly enhancing local estimation accuracy. By integrating parametric priors with nonparametric adjustments, the framework incorporates a tailored bandwidth selection strategy and naturally extends to nonparametric regression. Theoretical analysis and extensive simulations demonstrate that, even when the true density markedly departs from normality, the proposed estimator consistently outperforms conventional kernel density estimators across various Gaussian mixture models, particularly excelling in high-dimensional scenarios.
This study addresses the well-known boundary bias problem in existing nonparametric methods for estimating quantile density functions and their derivatives. To overcome this limitation, the authors propose a novel local polynomial-based estimation strategy that significantly enhances performance near boundaries. The proposed estimator demonstrates superior properties in terms of bias reduction, asymptotic variance, and boundary behavior compared to classical approaches. Moreover, the paper establishes the asymptotic normality of the estimator through rigorous theoretical analysis, thereby providing a more reliable foundation for nonparametric inference on both quantile densities and their derivatives. This advancement offers improved theoretical guarantees and practical utility for applications requiring accurate estimation in boundary regions.
This study addresses the challenges of non-negativity, normalization, and accuracy in estimating probability densities from empirical characteristic functions over a fixed time window. The authors propose a neural network approach trained directly in the Fourier domain, leveraging the closed-form characteristic function of a Gaussian–Laplace mixture model to enforce non-negativity and unit integral of the resulting density. The method is applicable to both independent and identically distributed data and dependent data resampling scenarios. As the first work to integrate neural networks with Fourier-domain training for density estimation, it derives a multidimensional L₂ error bound accounting for truncation, empirical, discretization, and sampling errors. Experiments demonstrate that the method matches the EM algorithm on Gaussian mixture benchmarks, substantially outperforms existing approaches on heavy-tailed distributions, exhibits L₂ error decay consistent with theoretical predictions, and successfully estimates the annual return distribution of the Australian stock market.
This study addresses the fundamental statistical problem of conditional density estimation by systematically comparing classical nonparametric approaches—such as single-index models, basis expansion methods (e.g., FlexCode), and DeepCDE—with modern generative models, including conditional GANs and conditional denoising diffusion probabilistic models. For the first time, these methods are evaluated within a unified and reproducible framework using metrics like mean squared error and Wasserstein distance to assess their accuracy, flexibility, and computational cost in estimating conditional means and standard deviations. The analysis clarifies the performance boundaries and practical applicability of each approach, offering both theoretical insights and actionable guidance for selecting appropriate methods in predictive modeling, uncertainty quantification, and probabilistic inference tasks.
This work addresses the computational expense and poor sample efficiency associated with modeling high-dimensional, non-Gaussian probability density functions in nonlinear dynamical systems. To overcome these challenges, the authors propose an efficient estimation approach based on a semi-nonparametric (SNP) density model. The method constructs a strictly positive density representation using Hermite polynomial basis functions and integrates Monte Carlo integration with a convex relaxation optimization strategy to significantly enhance the accuracy and stability of density and quantile estimation under limited sample sizes. Experimental results on the Lorenz chaotic system demonstrate that the proposed method accurately captures complex non-Gaussian structures and reliably computes quantiles using substantially fewer samples than conventional Monte Carlo techniques.