auxiliary gradient prediction

Design and evaluate surrogate predictors and sampling procedures that estimate per-example gradients or their summaries cheaply and use those auxiliary predictions to guide sample selection, weighting, or model-assisted/stratified sampling for stochastic optimization. Build models and analysis pipelines to quantify and reduce bias and variance of the resulting gradient estimators, integrate the predictions with optimizers via sample weights or selection without changing optimizer dynamics, and validate correlation with true gradients.

auxiliarygradientprediction

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.07
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the high variance in stochastic gradient estimation, which often leads to unstable convergence, slow training, and limited generalization in deep learning. To mitigate this issue, the paper introduces survey sampling theory into optimization for the first time, proposing a model-assisted sampling framework. Treating the dataset as a finite population, the method leverages an auxiliary gradient prediction model to construct low-variance gradient estimators, seamlessly integrating with momentum-based optimizers like AdamW without altering their dynamics. Empirical results demonstrate that the approach significantly improves performance in 71–86% of experimental settings across synthetic and six benchmark datasets, achieving superior generalization in approximately half the usual training time. The framework unifies and generalizes both uniform sampling and efficient, auxiliary-information-driven sampling strategies.

gradient estimationmini-batch samplingoptimization

This work addresses the challenge of risk prediction when acquiring true outcomes is prohibitively expensive, allowing labels for only a subset of samples. The authors propose a surrogate-assisted optimal sampling framework that, under a fixed annotation budget, leverages covariates, surrogate variables, and an initial estimator to construct a sampling strategy minimizing the expected out-of-sample cross-entropy loss. Coupled with an inverse probability weighted cross-entropy estimator for model training, this approach achieves—without requiring access to true responses during design—theoretically optimal sampling. It uniquely guarantees predictive optimality, robustness to surrogate misspecification, and stability in settings with rare outcomes. Both theoretical analysis and empirical experiments demonstrate that the method significantly outperforms existing approaches, particularly when the surrogate is imperfect or the event of interest is rare.

budget allocationmeasurement constraintsoptimal sampling

In many randomized trials, outcomes such as essays or open-ended responses must be manually scored as a preliminary step to impact analysis, a process that is costly and limiting. Model-assisted estimation offers a way to combine surrogate outcomes generated by machine learning or large language models with a human-coded subset, yet typical implementations use simple random sampling and therefore overlook systematic variation in surrogate prediction error. We extend this framework by incorporating stratified sampling to more efficiently allocate human coding effort. We derive the exact variance of the stratified model-assisted estimator, characterize conditions under which stratification improves precision, and identify a Neyman-type optimal allocation rule that oversamples strata with larger residual variance. We evaluate our methods through a comprehensive simulation study to assess finite-sample performance. Overall, we find stratification consistently improves efficiency when surrogate prediction errors exhibit structured bias or heteroskedasticity. We also present two empirical applications, one using data from an education RCT and one using a large observational corpus, to illustrate how these methods can be implemented in practice using ChatGPT-generated surrogate outcomes. Overall, this framework provides a practical design-based approach for leveraging surrogate outcomes and strategically allocating human coding effort to obtain unbiased estimates with greater efficiency. While motivated by text-as-data applications, the methodology applies broadly to any setting where outcome measurement is costly or prohibitive, and can be applied to comparisons across groups or estimating the mean of a single group.

human codingmodel-assisted estimationprediction error

Gradient-based Sample Selection for Faster Bayesian Optimization

Apr 10, 2025
QW
Qiyu Wei
🏛️ University of Manchester | National University of Singapore | Southern University of Science and Technology

Bayesian optimization (BO) suffers from poor scalability to large-budget settings due to the $O(n^3)$ time complexity of Gaussian process (GP) modeling. This work proposes a gradient-driven subset selection mechanism that significantly reduces GP fitting cost while preserving surrogate model fidelity. For the first time, gradient information is leveraged to jointly assess sample diversity and representativeness, enabling theoretically grounded sublinear regret—specifically, $O(sqrt{T}log T)$—thereby breaking the computational bottleneck inherent in full-data GP modeling. The method integrates GP regression, gradient-guided sampling, and subset optimization. Empirical evaluation on synthetic and real-world benchmarks demonstrates several-fold speedup in GP fitting time, while maintaining optimization performance comparable to standard BO. Thus, the approach achieves a favorable trade-off between computational efficiency and optimization effectiveness.

Maintaining optimization performance while lowering resource requirementsReducing computational cost of Gaussian process in Bayesian optimizationSelecting diverse samples using gradient information for efficiency

Two-stage Design for Failure Probability Estimation with Gaussian Process Surrogates

Oct 06, 2024
AS
Annie S. Booth
🏛️ Virginia Tech | Penn State

This work addresses the challenge of estimating small failure probabilities under stochastic inputs in computationally expensive deterministic simulations. We propose a two-stage adaptive budget allocation framework: in Stage I, a Gaussian process surrogate is sequentially trained using a contour-localization strategy; in Stage II, remaining simulation budget is greedily allocated to critical regions—guided by classification entropy—to perform high-fidelity evaluations. A hybrid Monte Carlo estimator is then constructed by integrating surrogate predictions with observed high-fidelity responses. Our method introduces the first “exploration–exploitation decoupled” budget allocation paradigm, overcoming reliability limitations inherent in pure surrogate-based Monte Carlo and importance sampling. Experiments across multiple benchmark functions and an airfoil flow simulation demonstrate that the approach achieves significantly improved accuracy and robustness using only several hundred high-fidelity evaluations.

Estimating failure probabilities with limited computational budgetImproving efficiency over existing sequential contour location methodsOptimizing surrogate model training for accurate classification

Latest Papers

What's happening recently
View more

This work addresses the high cost of ground-truth evaluation in chemical and materials design, where existing machine learning surrogate models often lack reliability guarantees. Departing from conventional reliance on prediction accuracy metrics such as R²—which can paradoxically increase the risk of worst-case selections—the study proposes “rank preservation” as a core criterion for surrogate validation. It formally introduces the concept of “selection tax” and derives its theoretical upper and lower bounds. A safety certification framework for surrogates is established through selection-aware auditing, rank correlation analysis, and multi-task ground-truth validation. Experiments demonstrate that the proposed audit statistics achieve Spearman correlations of 0.80–0.99 with actual search performance, substantially outperforming R² (as low as 0.33). Certified screening strategies based on this framework reduce evaluation costs by up to 25-fold.

experimental replacementmodel validationselection bias

High-dimensional stochastic agent-based models (ABMs) are notoriously difficult to analyze systematically due to the curse of dimensionality and inherent stochasticity. This work proposes a multi-stage automated exploration framework that first employs model-driven experimental design to identify key variables and partition the parameter space, then leverages machine learning surrogate models to efficiently capture residual nonlinear interaction effects. The approach operates without human intervention, automatically detecting unstable regions within the simulator and enabling robust sensitivity analysis and policy testing. Applied to a predator–prey case study, the framework successfully isolates dominant variables and highly sensitive nonlinear regimes, substantially enhancing the efficiency and reliability of ABM exploration.

Agent-Based Modelscurse of dimensionalitynonlinear interactions

This study addresses the low sample efficiency in expensive simulator inference and the underutilization of gradient information. We propose an active learning framework based on Bayesian optimization that integrates gradients obtained via automatic differentiation into Gaussian process surrogate models. Furthermore, this work presents the first systematic evaluation of the differential augmentation benefits between forward-mode and reverse-mode gradients under finite computational budgets. Experimental results demonstrate that reverse-mode gradients significantly accelerate convergence, with optimization gains sufficient to offset the additional computational overhead, whereas forward-mode gradients yield limited improvements. Overall, this research provides critical empirical evidence supporting the application of gradient-enhanced surrogate models for efficient simulation-based inference.

active learningBayesian inferenceexpensive simulators

Hot Scholars

FG

Feng Gao

Tsinghua University
Reinforcement LearningRobot Learning
MD

Mehdi Dastani

Professor, Chair for Intelligent Systems, Utrecht University
multi-agent systemsartificial intelligencecomputer science
LH

Liam Hodgkinson

University of Melbourne
probabilistic machine learningdeep learning theory
SW

Shihan Wang

Utrecht University
Machine LearningReinforcement LearningSocial Network Analysis