score function estimation

Designs and analyzes statistical estimators and algorithms that recover or approximate the score function (the gradient of the log-density) from samples, using approaches such as score matching, likelihood-ratio gradient estimation, empirical-risk-minimization with constraints, and constructions of unbiased score-function estimators. Works to prove error guarantees and minimax rates and to build function-class approximators (for example, neural networks with controlled complexity) that approximate scores on compact sets while explicitly managing estimator bias, variance, and dependence on ambient dimension.

scorefunctionestimation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.14
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the challenge of accurately estimating the score function of a probability measure from limited samples to enhance the generation quality of score-based generative models while mitigating overfitting. To this end, it introduces Sobolev space regularization into score function estimation for the first time, establishing a learning framework based on empirical risk minimization. The proposed method achieves minimax-optimal estimation rates on the flat torus and provides theoretical guarantees for score-based generative models, effectively balancing estimation accuracy with generalization capability.

empirical risk minimizationminimax ratesscore function estimation

Learning the score under shape constraints

Dec 16, 2025
RM
Rebecca M. Lewis
🏛️ University of Oxford | London School of Economics and Political Science | Nanjing University | Rutgers University | University of Cambridge

This paper investigates the minimax risk of score function estimation under log-concave distributions, focusing on the joint impact of tail behavior and smoothness on estimation accuracy. We introduce a novel density class that simultaneously imposes quantitative tail control and $(eta, L)$-Hölder smoothness constraints. For this class, we establish the optimal convergence rate—$n^{-eta/(2eta+1)}$, up to logarithmic factors, in the $L^2$ norm—for score function estimation—a first such result. Theoretically, this rate is strictly faster than those attainable under shape constraints alone or smoothness constraints alone when $eta in [1,2)$, revealing a synergistic gain from their combination. Methodologically, we construct a locally adaptive multiscale estimator based on uniform confidence bands, enabling automatic adaptation to heterogeneous smoothness across the domain. Our results unify and extend the theoretical frontiers of both shape-constrained estimation and nonparametric smooth function estimation.

Analyzes tail behavior impact and smoothness interplay in score estimation.Develops adaptive estimators for score functions with controlled growth and Hölder conditions.Estimates minimax risk for score estimation under log-concave shape constraints.

Minimax Semiparametric Learning With Approximate Sparsity

Dec 27, 2019
JB
Jelena Bradic
🏛️ Cornell University | MIT | Brandeis University

This paper addresses the optimal estimation of linear, mean-square-continuous functionals—such as average treatment effects and average derivatives—in high-dimensional approximately sparse regression. Existing methods rely on exact sparsity assumptions and thus fail to accommodate approximate sparsity arising from nonparametric basis expansions (e.g., splines, kernels, or polynomials); moreover, their minimax optimality remains unestablished. To bridge this gap, we propose a class of automated debiased machine learning estimators that require only single or weak cross-fitting. These estimators relax the convergence rate requirement on base learners to $n^{-1/4}$—substantially weaker than the conventional product-rate condition. Leveraging minimax theory and semiparametric efficiency bounds, we rigorously establish their $sqrt{n}$-consistency, asymptotic normality, and semiparametric efficiency. Under mild regularity conditions, they achieve the optimal convergence rate for the target functionals.

Developing minimax rates for approximately sparse semi-parametric modelsEstimating linear functionals in high-dimensional non-sparse modelsExploring root-n estimation without rate double robustness

On diffusion-based generative models and their error bounds: The log-concave case with full convergence estimates

Nov 22, 2023
SB
Stefano Bruno
🏛️ University of Edinburgh | The Hong Kong University of Science and Technology | Ulsan National Institute of Science and Technology | Imperial College London | The Alan Turing Institute | National Technical University of Athens

This work studies the convergence of diffusion models under strongly log-concave target distributions, establishing for the first time a joint upper bound on score estimation and sampling errors **without assuming Lipschitz continuity of the true score function**. Methodologically, it introduces an L² score estimation assumption based on the expected behavior of stochastic optimizers and constructs a novel auxiliary diffusion process; it further employs Lipschitz function approximation, probabilistic analysis, SDE theory, and optimization tools. Key contributions include: (1) achieving the current best-known Wasserstein-2 convergence rate; (2) deriving explicit, tight error bounds for sampling from unknown-mean Gaussians; and (3) obtaining dimensionally optimal upper bounds on convergence rates—substantially improving prior results and providing theoretical foundations for stochastic optimizer design.

Convergence of diffusion-based generative modelsError bounds in log-concave distributionsOptimization for score approximation sampling

Optimal convex $M$-estimation via score matching

Mar 25, 2024
OY
Oliver Y. Feng
🏛️ University of Bath | Rutgers University | University of Cambridge

For linear regression with unknown noise distribution, this paper proposes a data-driven method to construct convex loss functions such that the resulting empirical risk minimizer achieves the minimal asymptotic variance among all convex M-estimators. The key innovation is the first integration of score matching into the semiparametric convex M-estimation framework, unifying log-concave projection of the error density and minimization of Fisher divergence. For non-log-concave errors (e.g., Cauchy), the method automatically yields Huber-type optimal convex losses. We prove that the proposed estimator attains the minimal asymptotic covariance over all convex M-estimators; under Cauchy errors, its relative efficiency versus the oracle MLE exceeds 0.87. The algorithm combines nonparametric estimation of density derivatives, Fisher divergence optimization, and convex programming—ensuring both theoretical optimality and computational efficiency. Numerical experiments confirm its superior finite-sample performance.

Achieve robust, efficient estimation without log-concave noise assumptionsConstruct optimal convex loss for regression coefficient estimationDevelop nonparametric score matching for noise distribution approximation

Latest Papers

What's happening recently
View more

This work addresses the challenge of parameter inference and uncertainty quantification in generative models where the likelihood function is intractable. We propose a likelihood-free inference framework that approximates the gradient of the log-likelihood via structured score matching, enabling efficient estimation through gradient-based optimization combined with bootstrap resampling. The key innovation lies in a regularized neural network architecture that embeds statistical structure and a tailored score matching estimator, which together ensure theoretical convergence while substantially improving estimation accuracy and scalability. Numerical experiments demonstrate that the proposed method outperforms existing approaches in both parameter estimation accuracy and the reliability of uncertainty quantification.

generative modelsintractable likelihoodlikelihood-free inference

This work investigates whether the approximation capability of neural networks for score functions can guarantee that the distribution generated by the reverse diffusion process converges to the target data distribution. By integrating Hornik’s universal approximation theorem, Girsanov’s transformation on path space, and the data processing inequality for relative entropy, the study establishes—for the first time—an explicit quantitative relationship between the error in score function approximation and the discrepancy of the generated distribution. The resulting upper bound precisely links neural network approximation accuracy, noise scheduling, and prior mismatch, providing a rigorous KL divergence guarantee for score-based diffusion models grounded in classical approximation theory, without requiring any statistical assumptions.

distribution approximationKL divergenceneural network approximation

This study investigates the theoretical foundations of operator learning, with a focus on convergence rates and fundamental statistical limits. By integrating tools from statistical learning theory, approximation theory, and the framework of holomorphic operators, the work establishes a unified error analysis framework to systematically derive generalization error bounds for empirical risk minimization. Under generalized regularity conditions, it further establishes minimax-optimal statistical lower bounds, revealing an intrinsic trade-off between sample complexity and model approximation capacity. The analysis delineates current theoretical boundaries in operator learning and identifies several key open problems, offering new perspectives to guide future theoretical advances in the field.

convergence ratesminimax theoryoperator learning

This work addresses the challenge of high-dimensional Gaussian sampling, where conventional gradient-based oracles—providing only matrix-vector products with the precision matrix—lead to a sampling complexity that scales as √κ with the condition number κ. To overcome this limitation, the paper introduces a novel oracle based on smoothed scores, defined as the gradient of the log-density of a Gaussian-convolved distribution. By integrating geometrically spaced noise levels, sinc-quadrature rational approximation, coordinate quantization, and dithered post-processing, the proposed method eliminates the √κ dependence for the first time, reducing the condition-number dependence to logarithmic. Under total variation error δ_TV, the query complexity is O((log κ + log(e√d/δ_TV))·log(e√d/δ_TV)), and for fixed dimension and accuracy, the bit complexity achieves O(log²κ). An information-theoretic lower bound of Ω(log κ) is also established, demonstrating near-tightness.

condition numbergradient oraclequery complexity

Hot Scholars

DS

Dario Shariatian

PhD student, INRIA Paris
deep learninggenerative modelsdiffusion models
AD

Alain Durmus

Ecole polytechnique
Machine learningStatistics
SR

Sebastian Reich

University of Potsdam
Numerical AnalysisData AssimilationBayesian Inference
HY

Hong Ye Tan

Hedrick Assistant Adjunct Professor, UCLA
Machine LearningOptimizationInverse Problems
DZ

Delu Zeng

Professor with EE in South China University of Technology
Machine learningImage ProcessingBayesian LearningComputational Science