Variance-Optimal Control Variates for Learning with Black-box Feedback

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the suboptimal gradient estimation variance in black-box feedback learning, where conventional value functions used as control variates suffer from parameter sharing, and even exact value functions may remain far from the variance lower bound. To overcome this limitation, the work proposes an unbiased correction framework grounded in residual variance decomposition. It formally proves that the residual variance admits a cross-term-free decomposition with a closed-form optimal solution, which is practically realized through neural network parameter projection to achieve variance minimization. Both theoretical analysis and empirical evaluations demonstrate that the proposed approach significantly reduces residual variance and comprehensively improves learning efficiency across diverse tasks. The implementation code has been made publicly available.
📝 Abstract
Modern models increasingly learn through black-box oracles such as humans, optimization solvers, and external tools that provide feedback without exposing their internal mechanisms. A common remedy is to learn an (action-)value function as a control variate. In this paper, we first observe that even an exact action-value function can be arbitrarily far from variance-optimal. We show that this gap arises because the value function minimizes the noise in each action's own gradient term, while an action can still affect the rest of the gradient estimator through shared parameters. A simple unbiased correction, at no extra oracle cost, can still reduce its variance by an arbitrarily large factor. Motivated by this, we then prove that the residual variance can be decomposed exactly by actions with no cross terms. This decomposition yields a closed-form variance-minimizing correction for neural-network parameters, which can be computed by a simple projection. Empirically, our correction consistently reduces the variance left by the value function and improves learning across all tasks. The source code for all experiments is available at https://github.com/Zihao-Kevin/black_box_opt.
Problem

Research questions and friction points this paper is trying to address.

black-box feedback
variance reduction
control variates
action-value function
gradient estimation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Black-box Feedback
Control Variates
Variance Reduction
Action-Value Function
Residual Variance Decomposition
🔎 Similar Papers
Z
Zihao Zhao
School of Computational Science and Engineering, Georgia Institute of Technology
S
Shuhan Zhang
School of Computational Science and Engineering, Georgia Institute of Technology
Kai Wang
Kai Wang
Harbin Institute of Technology at Weihai Campus, Professor
Trustworthy AINetwork Intrusion Detection