dual distribution estimation

Designs and implements methods to estimate and model dual probability distributions of data (for example class-wise or positive/negative components) using parametric or nonparametric density models. Builds algorithms that leverage those distribution estimates for distribution-based NTTA, online out-of-distribution filtering, and training-free, scalable inference.

dualdistributionestimation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.3
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the fundamental statistical problem of conditional density estimation by systematically comparing classical nonparametric approaches—such as single-index models, basis expansion methods (e.g., FlexCode), and DeepCDE—with modern generative models, including conditional GANs and conditional denoising diffusion probabilistic models. For the first time, these methods are evaluated within a unified and reproducible framework using metrics like mean squared error and Wasserstein distance to assess their accuracy, flexibility, and computational cost in estimating conditional means and standard deviations. The analysis clarifies the performance boundaries and practical applicability of each approach, offering both theoretical insights and actionable guidance for selecting appropriate methods in predictive modeling, uncertainty quantification, and probabilistic inference tasks.

conditional distribution estimationgenerative modelsnonparametric methods

This work addresses the challenge that traditional discrete probability distributions rely on manually derived analytical forms, hindering the automatic discovery of interpretable models. We propose Symbolic Density Estimation (SDE), a novel framework that, for the first time, integrates structural priors, evolutionary search, and validity-aware parameter inference to automatically discover closed-form probability mass functions within a structured symbolic space composed of elementary mathematical operations. SDE accommodates complex distributional features such as zero-inflation and finite mixtures. We introduce the first systematic benchmark dataset for this task and demonstrate that SDE accurately recovers all target distribution families. On real-world data, SDE discovers concise, interpretable mixture models that achieve superior goodness-of-fit compared to standard methods.

discrete distributionsinterpretable modelsprobability mass functions

Transforming Conditional Density Estimation Into a Single Nonparametric Regression Task

Nov 23, 2025
AG
Alexander G. Reisach
🏛️ CNRS | Université Paris Cité | MODAL'X | Université Paris Nanterre | Harvard University

This paper addresses the challenge of applying powerful regression models directly to high-dimensional conditional density estimation (CDE). We propose a novel framework that reformulates CDE as a single nonparametric regression task by constructing labeled auxiliary samples, thereby mapping density estimation into a standard regression problem. Our approach imposes no parametric assumptions on the conditional density and enables plug-and-play integration of arbitrary state-of-the-art regressors—including deep neural networks and gradient-boosted trees. We establish theoretical consistency: the proposed estimator converges almost surely to the true conditional density as the sample size tends to infinity. Extensive experiments on synthetic data, U.S. Census records, and satellite imagery demonstrate that our method significantly outperforms existing CDE benchmarks across diverse real-world scenarios. Moreover, the results align with domain-specific prior knowledge, underscoring the method’s balance of theoretical rigor and practical efficacy.

Achieves state-of-the-art performance on survey and satellite datasetsLeverages neural networks for high-dimensional density estimationTransforms conditional density estimation into nonparametric regression

Easy Conditioning far beyond Gaussian

Sep 24, 2024
AF
Antoine Faul
🏛️ University of Bern

This work addresses the longstanding limitation in conditional density estimation—namely, the absence of closed-form solutions for multivariate conditional densities under non-Gaussian assumptions. We propose a generative conditional density estimation framework grounded in copula modeling and analytic conditionalization in latent space. Methodologically, we first establish the inheritability of “conditional stability” under mixture and transformation operations, thereby extending analytically tractable conditional families to non-Gaussian, nonlinear, and cross-dimensional settings. The core components include a Gaussian Mixture Copula Model (GMCM), an explicit latent-space conditionalization mechanism, and joint copula modeling. Experiments on synthetic and real-world datasets demonstrate substantial improvements in conditional density estimation accuracy and robustness to missing data imputation. Crucially, our approach enables efficient, differentiable, and sampling-free deterministic conditional inference.

Applying copula-based models for density estimation and imputationDeveloping generative method for estimating conditional distributionsExtending analytical conditioning beyond Gaussian distributions

A Nonparametric Maximum Likelihood Approach to Mixture of Regression

Aug 22, 2021
HJ
Hansheng Jiang
🏛️ University of Toronto | University of California, Berkeley

This paper addresses modeling population heterogeneity in mixed linear regression (i.e., random-coefficient) models. We propose a fully nonparametric maximum likelihood estimator (NPMLE) for the unknown mixing distribution (G^*), without prespecifying its parametric form or the number of components. Our method directly computes the NPMLE via convex optimization—yielding the first rigorous proof of its existence—and establishes, for finite samples, an optimal parametric-rate bound (up to logarithmic factors) on the Hellinger estimation error, circumventing discretization-induced bias inherent in conventional approaches. Theoretically and empirically, the estimator achieves both statistical efficiency and computational tractability: it significantly outperforms EM-based parametric methods on both discrete and continuous mixture simulations, as well as two real-world datasets, demonstrating strong robustness and practical utility.

Achieves near-parametric rates in conditional density estimationEstimates unknown distribution of regression coefficients nonparametricallyProvides posterior-based individualized coefficient inference empirically

Latest Papers

What's happening recently
View more

This work proposes a semiparametric density estimation approach that addresses the inefficiency of traditional kernel density estimation in high-dimensional settings or when the underlying distribution deviates substantially from assumed parametric forms. The method multiplies a parametric initial guess—such as a normal distribution—by a nonparametric kernel-based correction factor, thereby preserving robustness against departures from the parametric family while significantly enhancing local estimation accuracy. By integrating parametric priors with nonparametric adjustments, the framework incorporates a tailored bandwidth selection strategy and naturally extends to nonparametric regression. Theoretical analysis and extensive simulations demonstrate that, even when the true density markedly departs from normality, the proposed estimator consistently outperforms conventional kernel density estimators across various Gaussian mixture models, particularly excelling in high-dimensional scenarios.

density estimationkernel estimatornonparametric

This study addresses the challenge of enhancing density estimation performance by integrating the structural advantages of parametric models while preserving the flexibility of nonparametric methods. The authors propose a locally parametrized nonparametric density estimator that, for each point \(x\), estimates an optimal local parameter \(\hat{\theta}(x)\) via kernel-smoothed likelihood, yielding an estimator of the form \(f(x, \hat{\theta}(x))\). This approach achieves near-full-likelihood efficiency under correct model specification and retains nonparametric robustness under misspecification, effectively serving as a semiparametric realization of higher-order kernel methods. Theoretical analysis and empirical experiments demonstrate that the proposed estimator exhibits variance comparable to classical kernel density estimation but substantially reduced bias, leading to significantly improved accuracy in neighborhoods of well-specified parametric models.

density estimationlocal likelihoodnonparametric

This work addresses the computational expense and poor sample efficiency associated with modeling high-dimensional, non-Gaussian probability density functions in nonlinear dynamical systems. To overcome these challenges, the authors propose an efficient estimation approach based on a semi-nonparametric (SNP) density model. The method constructs a strictly positive density representation using Hermite polynomial basis functions and integrates Monte Carlo integration with a convex relaxation optimization strategy to significantly enhance the accuracy and stability of density and quantile estimation under limited sample sizes. Experimental results on the Lorenz chaotic system demonstrate that the proposed method accurately captures complex non-Gaussian structures and reliably computes quantiles using substantially fewer samples than conventional Monte Carlo techniques.

chaotic systemsdata efficiencynon-Gaussian density estimation

This work addresses the high computational cost and lack of convergence rate guarantees associated with nonparametric maximum likelihood estimation (NPMLE) in exponential family mixture models. The authors propose a data-compression-based acceleration strategy that, for the first time, reduces the likelihood evaluation complexity of NPMLE to logarithmic order. They establish rigorous statistical theory for the resulting approximate estimator, demonstrating that the proposed method achieves near-parametric convergence rates for marginal density estimation while substantially lowering computational overhead.

computation efficiencyconvergence rateexponential family mixtures

This study addresses the challenge of nonparametric density estimation for univariate grouped data that only provide interval frequencies. It proposes a novel method—Mean-Adjusted Log-Concave (MALC) estimation—that requires neither prior distributional assumptions nor bandwidth selection. By formulating an optimization framework under log-concavity constraints, MALC integrates both interval frequencies and interval means to reconstruct the underlying density function. As the first approach for grouped data that avoids reliance on kernel functions or parametric assumptions, MALC demonstrates superior robustness and estimation accuracy across a wide range of simulation scenarios, including varying distribution shapes, sample sizes, and bin widths.

bandwidth-freedensity estimationgrouped data

Hot Scholars

IH

Ioannis Havoutis

Associate Professor, Oxford Robotics Institute, University of Oxford
RoboticsMachine LearningLegged Robotics
AR

Andrew Rosemberg

Optimisation Researcher
Operations ResearchOptimizationMachine LearningEnergy