Score
Designs and implements neural conditional density estimators that approximate Bayesian posteriors concentrated around a specific observed dataset by restricting training and inference to regions near that observation. Develops and evaluates localization mechanisms (distance-aware weighting, local loss functions, or predictive-focused network architectures) and trains/optimizes these estimators for calibrated predictive scoring and robustness to model misspecification.
Existing neural density estimation methods exhibit strong empirical performance but struggle to simultaneously enforce non-negativity and unit-mass constraints, and lack theoretical guarantees for adaptive convergence on low-dimensional structured densities. This paper proposes a structure-agnostic, classification-induced neural density estimation framework: by reformulating density ratio estimation as a binary classification task, it implicitly satisfies probabilistic constraints without explicit regularization. We establish theoretical guarantees showing that, under high-dimensional densities supported on low-dimensional manifolds, the estimator achieves adaptive convergence rates strictly faster than standard nonparametric rates. Moreover, the framework naturally yields an efficient sampling mechanism. Extensive experiments on synthetic and real-world datasets demonstrate significant improvements in convergence speed, density estimation accuracy, and sample quality. To our knowledge, this work provides the first rigorous theoretical connection between the classification paradigm and neural density estimation.
This paper addresses the challenge of applying powerful regression models directly to high-dimensional conditional density estimation (CDE). We propose a novel framework that reformulates CDE as a single nonparametric regression task by constructing labeled auxiliary samples, thereby mapping density estimation into a standard regression problem. Our approach imposes no parametric assumptions on the conditional density and enables plug-and-play integration of arbitrary state-of-the-art regressors—including deep neural networks and gradient-boosted trees. We establish theoretical consistency: the proposed estimator converges almost surely to the true conditional density as the sample size tends to infinity. Extensive experiments on synthetic data, U.S. Census records, and satellite imagery demonstrate that our method significantly outperforms existing CDE benchmarks across diverse real-world scenarios. Moreover, the results align with domain-specific prior knowledge, underscoring the method’s balance of theoretical rigor and practical efficacy.
For high-dimensional simulator-based models with intractable likelihoods, this paper proposes an efficient and stable Sequential Neural Posterior Estimation (SNPE) method. The approach employs conditional neural density estimation within a sequential simulation framework, augmented by an adaptive calibration kernel mechanism—novelly introduced herein—to dynamically adjust kernel weights during inference. To further enhance stability and accelerate convergence, we integrate importance-weighted gradient variance reduction with Monte Carlo loss optimization. This combination effectively mitigates the inference bottlenecks inherent in high-dimensional settings while preserving posterior approximation accuracy. Extensive experiments on multiple benchmark simulators and real-world high-dimensional datasets demonstrate that our method achieves over a two-fold speedup in training time and reduces posterior approximation error by more than 30% compared to standard SNPE and other state-of-the-art approaches.
Deep learning models often suffer from miscalibrated predictive uncertainty: their predicted confidence intervals exhibit coverage rates substantially below nominal levels (e.g., 90% intervals covering <90% of test outcomes) and degraded sharpness. This work proposes a probabilistic recalibration training paradigm based on low-dimensional density estimation, jointly improving calibration and sharpness without compromising overall model performance. It is the first general-purpose distributional calibration method applicable to arbitrary models—including deep neural networks—with theoretical guarantees and a consistent convergence bound. The approach integrates maximum-likelihood enhancement with Bayesian modeling principles, yielding significant improvements in both calibration and predictive accuracy for linear and deep Bayesian models; notably, 90% confidence intervals achieve coverage rates approaching the nominal level. The implementation, including code and tutorials, is publicly available.
To address the scalability challenges of Bayesian learning under big data and large models—stemming from high-dimensional posterior approximation—this paper proposes a scalable Bayesian inference framework. Methodologically, it introduces a novel tempered stochastic gradient MCMC perspective, theoretically establishing the asymptotic unbiasedness of deep ensembles. It further provides the first systematic empirical validation of the cold posterior effect in large language models (LLMs), demonstrating improved uncertainty calibration and robustness via Bayesian approximation. Finally, it develops Posteriors, an open-source PyTorch library implementing a unified optimization-and-sampling paradigm, enabling efficient Bayesian inference for models with up to thousands of layers. Experiments across multiple benchmarks and LLM tasks show significant gains in predictive uncertainty calibration and out-of-distribution robustness.
This study addresses the limitations of existing neural density estimation methods, which rely on invertible network architectures and computationally expensive Jacobian determinant calculations, thereby constraining model flexibility and scalability. To overcome these challenges, this work proposes a Jacobian-free density estimation framework grounded in Bayesian generative modeling. By employing variational inference to approximate latent variable posteriors for constructing adaptive proposal distributions, and integrating bridge sampling techniques for precise density estimation, the approach directly transforms generative models into flexible density estimators. A key contribution is that this method entirely eliminates the dependence on invertible networks. Experimental results demonstrate significant improvements in both density estimation accuracy and structural recovery on synthetic datasets, while also exhibiting superior performance in anomaly detection tasks across real-world scenarios.
This work addresses the challenge of Bayesian inference in misspecified time series models when the likelihood function is intractable but data can be simulated. It proposes an adaptive distance learning approach grounded in scoring rule optimization, which unifies the frameworks of Approximate Bayesian Computation (ABC) and localized Neural Posterior Estimation with Predictive Function Networks (NPE-PFN). The method establishes a theoretical equivalence between adaptive distance learning and the estimation of optimal linear combination weights for predictors. Empirical evaluations on both synthetic and real-world datasets demonstrate that the proposed approach substantially improves posterior estimation accuracy and predictive performance, offering a novel and effective strategy for inference under model misspecification in dynamic settings.
This work proposes a nonparametric inference method for conditional functionals—such as the conditional mean—in settings where labeled data are scarce, unlabeled covariates are abundant, and a black-box predictor is available. The approach avoids parametric modeling assumptions by leveraging a data-adaptive kernel localization and a prediction-correction decomposition, which transforms conditional moment estimation into a weighted unconditional moment problem while incorporating the black-box predictor to reduce variance. Theoretical analysis establishes non-asymptotic error bounds, minimax optimal convergence rates, and asymptotic normality, along with an explicit variance decomposition that quantifies the contributions of both the predictor and the unlabeled data. Experiments demonstrate that the resulting confidence intervals achieve accurate coverage and are significantly narrower, marking the first method in a fully nonparametric framework to simultaneously attain validity and efficiency gains.
This study addresses the unclear optimal allocation of parameter stochasticity in Bayesian neural networks, highlighting the need to balance approximation capacity with computational efficiency. We propose a prior-scale-based deep weight decomposition method that achieves stochasticity sparsification by fitting Gaussian process priors via maximum mean discrepancy and employing a threshold-switching mechanism to convert low-scale parameters into deterministic optimization. Furthermore, we introduce a general density approximation certificate verifiable in linear time, revealing that hybrid schemes fundamentally constitute Type-II MAP stochastic approximation. Experiments demonstrate that the proposed approach attains fully stochastic network performance on UCI benchmarks using only half the deterministic parameters, while significantly outperforming existing baselines on bimodal tasks.
This study investigates the construction of Bayesian predictive inference methods with favorable asymptotic properties and offers a novel Bayesian interpretation of kernel density estimation. Focusing on two classes of predictive rules—classical kernel density estimators and their recursive variants—the work systematically examines their weak almost sure convergence under sequential observations by integrating nonparametric kernel methods, stochastic process theory, and weak convergence analysis. The analysis reveals that the classical estimator converges weakly almost surely to a probability measure with compact support, whereas its recursive counterpart converges to one with non-compact support. Beyond establishing weak almost sure convergence for both schemes, this research extends the theoretical foundations of kernel methods within Bayesian predictive inference and provides a fresh Bayesian perspective on kernel density estimation.