Score
Designs and analyzes statistical estimators and algorithms that recover or approximate the score function (the gradient of the log-density) from samples, using approaches such as score matching, likelihood-ratio gradient estimation, empirical-risk-minimization with constraints, and constructions of unbiased score-function estimators. Works to prove error guarantees and minimax rates and to build function-class approximators (for example, neural networks with controlled complexity) that approximate scores on compact sets while explicitly managing estimator bias, variance, and dependence on ambient dimension.
This work addresses the challenge of accurately estimating the score function of a probability measure from limited samples to enhance the generation quality of score-based generative models while mitigating overfitting. To this end, it introduces Sobolev space regularization into score function estimation for the first time, establishing a learning framework based on empirical risk minimization. The proposed method achieves minimax-optimal estimation rates on the flat torus and provides theoretical guarantees for score-based generative models, effectively balancing estimation accuracy with generalization capability.
This paper investigates the minimax risk of score function estimation under log-concave distributions, focusing on the joint impact of tail behavior and smoothness on estimation accuracy. We introduce a novel density class that simultaneously imposes quantitative tail control and $(eta, L)$-Hölder smoothness constraints. For this class, we establish the optimal convergence rate—$n^{-eta/(2eta+1)}$, up to logarithmic factors, in the $L^2$ norm—for score function estimation—a first such result. Theoretically, this rate is strictly faster than those attainable under shape constraints alone or smoothness constraints alone when $eta in [1,2)$, revealing a synergistic gain from their combination. Methodologically, we construct a locally adaptive multiscale estimator based on uniform confidence bands, enabling automatic adaptation to heterogeneous smoothness across the domain. Our results unify and extend the theoretical frontiers of both shape-constrained estimation and nonparametric smooth function estimation.
This paper addresses the optimal estimation of linear, mean-square-continuous functionals—such as average treatment effects and average derivatives—in high-dimensional approximately sparse regression. Existing methods rely on exact sparsity assumptions and thus fail to accommodate approximate sparsity arising from nonparametric basis expansions (e.g., splines, kernels, or polynomials); moreover, their minimax optimality remains unestablished. To bridge this gap, we propose a class of automated debiased machine learning estimators that require only single or weak cross-fitting. These estimators relax the convergence rate requirement on base learners to $n^{-1/4}$—substantially weaker than the conventional product-rate condition. Leveraging minimax theory and semiparametric efficiency bounds, we rigorously establish their $sqrt{n}$-consistency, asymptotic normality, and semiparametric efficiency. Under mild regularity conditions, they achieve the optimal convergence rate for the target functionals.
This work studies the convergence of diffusion models under strongly log-concave target distributions, establishing for the first time a joint upper bound on score estimation and sampling errors **without assuming Lipschitz continuity of the true score function**. Methodologically, it introduces an L² score estimation assumption based on the expected behavior of stochastic optimizers and constructs a novel auxiliary diffusion process; it further employs Lipschitz function approximation, probabilistic analysis, SDE theory, and optimization tools. Key contributions include: (1) achieving the current best-known Wasserstein-2 convergence rate; (2) deriving explicit, tight error bounds for sampling from unknown-mean Gaussians; and (3) obtaining dimensionally optimal upper bounds on convergence rates—substantially improving prior results and providing theoretical foundations for stochastic optimizer design.
For linear regression with unknown noise distribution, this paper proposes a data-driven method to construct convex loss functions such that the resulting empirical risk minimizer achieves the minimal asymptotic variance among all convex M-estimators. The key innovation is the first integration of score matching into the semiparametric convex M-estimation framework, unifying log-concave projection of the error density and minimization of Fisher divergence. For non-log-concave errors (e.g., Cauchy), the method automatically yields Huber-type optimal convex losses. We prove that the proposed estimator attains the minimal asymptotic covariance over all convex M-estimators; under Cauchy errors, its relative efficiency versus the oracle MLE exceeds 0.87. The algorithm combines nonparametric estimation of density derivatives, Fisher divergence optimization, and convex programming—ensuring both theoretical optimality and computational efficiency. Numerical experiments confirm its superior finite-sample performance.
This work addresses the challenge of parameter inference and uncertainty quantification in generative models where the likelihood function is intractable. We propose a likelihood-free inference framework that approximates the gradient of the log-likelihood via structured score matching, enabling efficient estimation through gradient-based optimization combined with bootstrap resampling. The key innovation lies in a regularized neural network architecture that embeds statistical structure and a tailored score matching estimator, which together ensure theoretical convergence while substantially improving estimation accuracy and scalability. Numerical experiments demonstrate that the proposed method outperforms existing approaches in both parameter estimation accuracy and the reliability of uncertainty quantification.
This work investigates whether the approximation capability of neural networks for score functions can guarantee that the distribution generated by the reverse diffusion process converges to the target data distribution. By integrating Hornik’s universal approximation theorem, Girsanov’s transformation on path space, and the data processing inequality for relative entropy, the study establishes—for the first time—an explicit quantitative relationship between the error in score function approximation and the discrepancy of the generated distribution. The resulting upper bound precisely links neural network approximation accuracy, noise scheduling, and prior mismatch, providing a rigorous KL divergence guarantee for score-based diffusion models grounded in classical approximation theory, without requiring any statistical assumptions.
This study investigates the theoretical foundations of operator learning, with a focus on convergence rates and fundamental statistical limits. By integrating tools from statistical learning theory, approximation theory, and the framework of holomorphic operators, the work establishes a unified error analysis framework to systematically derive generalization error bounds for empirical risk minimization. Under generalized regularity conditions, it further establishes minimax-optimal statistical lower bounds, revealing an intrinsic trade-off between sample complexity and model approximation capacity. The analysis delineates current theoretical boundaries in operator learning and identifies several key open problems, offering new perspectives to guide future theoretical advances in the field.
This work addresses the challenge of high-dimensional Gaussian sampling, where conventional gradient-based oracles—providing only matrix-vector products with the precision matrix—lead to a sampling complexity that scales as √κ with the condition number κ. To overcome this limitation, the paper introduces a novel oracle based on smoothed scores, defined as the gradient of the log-density of a Gaussian-convolved distribution. By integrating geometrically spaced noise levels, sinc-quadrature rational approximation, coordinate quantization, and dithered post-processing, the proposed method eliminates the √κ dependence for the first time, reducing the condition-number dependence to logarithmic. Under total variation error δ_TV, the query complexity is O((log κ + log(e√d/δ_TV))·log(e√d/δ_TV)), and for fixed dimension and accuracy, the bit complexity achieves O(log²κ). An information-theoretic lower bound of Ω(log κ) is also established, demonstrating near-tightness.