Extending Pathwise Gradients to Discrete Random Variables via Finite-Order Relaxation

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of high variance and hyperparameter sensitivity in pathwise gradient estimation for discrete variables by proposing a finite-order relaxation framework. Methodologically, we construct an exact pathwise gradient estimator that preserves hard-sample properties and introduce the first minimum-norm unbiased estimator. By integrating polynomial constrained optimization with probability distribution theory, we derive non-asymptotic bias bounds. Notably, this framework significantly reduces weight variance without requiring temperature tuning. Experimental results demonstrate that our approach matches or surpasses existing baselines across diverse model architectures while achieving superior computational efficiency throughout.
📝 Abstract
Pathwise gradients are preferred for continuous random variables because they are unbiased, low variance, and work with a single sample. For discrete variables, however, the pathwise identity cannot generally be exact for every differentiable function. We propose a general framework to construct finite-order exact pathwise gradient estimators for a range of common discrete variables such as Poisson. The estimator is the least-norm solution among all solutions that are unbiased for polynomials of degree at most. The resulting estimators preserve the hard forward sample, require no temperature tuning, and can be implemented in a few lines of codes. Against other admissible solutions, our estimator is unique and minimizes weight variance; in contrast, prior works use categorical variables or augmented representations to approximate non-categorical variables that induces excess variance and computations. To understand approximation bias for functions beyond the prescribed class, we also derive a non-asymptotic bias bound. In experiments our low order methods match or improve tuned baselines across linear, nonlinear and hierarchical latent-variable models, while out-speeding competitors in every runtime benchmark.
Problem

Research questions and friction points this paper is trying to address.

pathwise gradients
discrete random variables
gradient estimation
variance reduction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Pathwise Gradients
Discrete Random Variables
Finite-Order Relaxation
Gradient Estimators
Variance Minimization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.