Score
Designs and implements neural-network–based inference systems and gradient-based optimization procedures that estimate or optimize model parameters, latent variables, or hyperparameters by learning surrogates, likelihoods, or amortized posteriors. Builds differentiable surrogate models, inference networks, learned discrepancy/metric functions, and training pipelines used to calibrate simulators, tune generator hyperparameters, and compare or select candidate generative models via gradients computed through learned components.
This work addresses the challenge of parameter inference in stochastic processes, where conventional simulation-based inference methods suffer from computationally expensive likelihood evaluations and struggle to balance surrogate model accuracy against simulation cost under limited data. Breaking from the black-box assumption, this study introduces a novel approach that incorporates exact score information and a loss-gradient–based adaptive weighting mechanism into a probabilistic classification framework for neural likelihood surrogates, optimizing binary cross-entropy loss. The proposed method substantially enhances both the efficiency and accuracy of the surrogate model. Across multiple stochastic process benchmarks, it achieves downstream inference performance equivalent to using ten times more training data while requiring only 1.1× the original training time, effectively alleviating the data–cost trade-off bottleneck.
For stochastic simulation models with intractable likelihoods, existing score estimators based on noisy Monte Carlo ratio estimators suffer from bias and instability. Method: We propose the first gradient-based simulation parameter estimation framework, which eliminates ratio bias via a multi-timescale stochastic approximation algorithm, incorporates a nested simulation optimization architecture, and extends— for the first time—to neural network training. The method integrates stochastic approximation, multiscale optimization, nested Monte Carlo estimation, and asymptotic statistical analysis. Contributions/Results: We rigorously establish strong consistency, asymptotic normality, optimal convergence rate, and an optimal budget allocation strategy for the estimator. Numerical experiments demonstrate substantial improvements in estimation accuracy and significant reductions in computational cost.
This paper addresses the inefficiency and lack of scalability of manual hyperparameter tuning in large-scale machine learning. It systematically surveys hyperparameter optimization (HPO), unifying and classifying five mainstream paradigms: random/low-discrepancy search, bandit-based methods, Bayesian optimization, population-based (evolutionary) algorithms, and gradient-based differentiable optimization. The survey further extends to emerging settings—including online HPO, constrained HPO, and multi-objective HPO. Crucially, the work establishes novel theoretical connections between HPO and meta-learning as well as neural architecture search, yielding a comprehensive knowledge framework that articulates methodological principles, applicability boundaries, and inherent limitations. By clarifying the technical evolution and identifying key open challenges, this study provides a theoretically grounded yet practically actionable foundation for automated machine learning.
This work addresses the challenges of poor convergence and limited scalability in variational inference for Bayesian neural networks (BNNs). We propose a practical training framework that optimizes the evidence lower bound (ELBO) using stochastic gradient descent (SGD). For the first time, we systematically establish the theoretical feasibility of SGD for variational learning in BNNs and design a robust gradient estimation strategy to enable parameterized posterior approximation. By avoiding traditional complex sampling schemes or second-order optimization methods, our approach significantly reduces computational overhead. Experiments on five UCI regression benchmarks demonstrate that our method outperforms state-of-the-art BNN approaches in both root mean squared error (RMSE) and negative log-likelihood (NLL), thereby improving the accuracy and practicality of uncertainty quantification.
For high-dimensional simulator-based models with intractable likelihoods, this paper proposes an efficient and stable Sequential Neural Posterior Estimation (SNPE) method. The approach employs conditional neural density estimation within a sequential simulation framework, augmented by an adaptive calibration kernel mechanism—novelly introduced herein—to dynamically adjust kernel weights during inference. To further enhance stability and accelerate convergence, we integrate importance-weighted gradient variance reduction with Monte Carlo loss optimization. This combination effectively mitigates the inference bottlenecks inherent in high-dimensional settings while preserving posterior approximation accuracy. Extensive experiments on multiple benchmark simulators and real-world high-dimensional datasets demonstrate that our method achieves over a two-fold speedup in training time and reduces posterior approximation error by more than 30% compared to standard SNPE and other state-of-the-art approaches.
This study addresses the low sample efficiency in expensive simulator inference and the underutilization of gradient information. We propose an active learning framework based on Bayesian optimization that integrates gradients obtained via automatic differentiation into Gaussian process surrogate models. Furthermore, this work presents the first systematic evaluation of the differential augmentation benefits between forward-mode and reverse-mode gradients under finite computational budgets. Experimental results demonstrate that reverse-mode gradients significantly accelerate convergence, with optimization gains sufficient to offset the additional computational overhead, whereas forward-mode gradients yield limited improvements. Overall, this research provides critical empirical evidence supporting the application of gradient-enhanced surrogate models for efficient simulation-based inference.
This work addresses the inefficiency of global optimization when standard neural network surrogates are embedded into mixed-integer linear programs (MILPs), a challenge stemming from the lack of control over their structural properties. The authors propose a novel differentiable regularizer that, for the first time, approximates the full gradient of the LP relaxation gap with respect to network parameters, enabling direct optimization of key structural attributes such as big-M constants, the number of unstable neurons, and the LP relaxation gap itself. Built upon ReLU networks and MILP formulations, the method leverages gradients from LP dual variables and requires no custom automatic differentiation. Experiments demonstrate up to four orders of magnitude reduction in MILP solve time on nonconvex benchmark functions and two-stage stochastic programming problems, all while preserving predictive accuracy.
This work addresses the high computational cost of Monte Carlo estimation in Bayesian inverse problems, which often arises from large variances in quantities of interest under the posterior distribution. To mitigate this, the authors propose a conditional neural control variate method that learns generalizable control variates from joint samples of parameters and data to effectively reduce variance. The approach leverages a scalable neural architecture grounded in Stein’s identity, incorporating hierarchical coupling layers and tractable Jacobian trace computation, enabling generalization across different observational data without retraining. The required posterior score function can be derived from physical models, neural operators, or conditional normalizing flows. Demonstrated on a Darcy flow inverse problem, the method achieves substantial variance reduction even when using learned score approximations in place of analytical scores.
This work proposes a unified framework that integrates hierarchical Bayesian inference with data-driven closure learning to address the inverse problem of model calibration in multiphysics systems, where unknown parameters and incomplete dynamical laws pose significant challenges. The approach jointly infers system-specific parameters and shared unknown dynamics across multiple related systems through a hierarchical structure. Neural networks—such as Fourier Neural Operators (FNOs) and parameterized Physics-Informed Neural Networks (PINNs)—are employed to construct closure models for ODEs/PDEs. Efficient posterior inference is achieved via maximum marginal likelihood estimation combined with ensemble Metropolis-adjusted Langevin algorithm (MALA) sampling. Furthermore, an adaptive surrogate forward model and a bilevel optimization strategy are introduced to substantially reduce the computational cost associated with repeated forward solves. Experiments demonstrate that the framework achieves high calibration accuracy while enabling efficient joint modeling and computational acceleration across systems.
This work proposes a novel approach to simulation-based inference by integrating large language model–driven program synthesis, enabling joint inference of both model structure and parameters—a capability lacking in traditional methods that rely on fixed, pre-specified simulator architectures. By automatically generating and iteratively refining candidate simulator programs from natural language descriptions, the method transcends rigid modeling assumptions and facilitates the discovery of plausible models directly from open-ended prompts. Empirical evaluations across diverse domains—including deterministic dynamical systems, stochastic epidemic models, and gravitational lensing image analysis—demonstrate its ability to accurately identify data-supported model families, revealing the interplay between the informativeness of observed data and the identifiability of candidate models.