Score
Design and implement training procedures that enable backpropagation through nondifferentiable discrete operations by replacing true gradients with straight-through estimator (STE) approximations, including custom forward/backward passes for binary, quantized, or hard-selection units. Analyze and mitigate STE-induced gradient bias and variance, select or adapt STE variants for particular discrete components, and validate their impact on optimization stability and end-to-end model performance.
Straight-through estimators (STE) are widely used in quantized training to approximate gradients through non-differentiable discrete operations, yet their finite-sample theoretical guarantees have remained elusive. Method: Leveraging tools from compressed sensing and dynamical systems theory, we develop a finite-sample convergence analysis for STE in quantized neural network training. Contribution/Results: We derive an explicit upper bound on the sample complexity required for global convergence—scaling with data dimension—and rigorously prove that STE converges to the global optimum under finite samples. Moreover, we uncover a novel oscillatory behavior in the presence of label noise: STE iterates periodically escape from and return to the optimal solution. Our results establish sample size as a critical determinant of STE’s success and provide the first theoretically grounded framework for analyzing gradient approximation in quantized training.
Straight-through estimators (STE) are widely adopted in quantized neural network training due to the discrete, non-differentiable nature of the objective, yet the mechanistic role of hyperparameters—such as bit-width and quantization range—in governing STE’s learning dynamics remains poorly understood. Method: We rigorously derive, in the high-dimensional limit, that STE dynamics converge to a deterministic differential equation. We analyze fixed points to quantify the asymptotic bias relative to unquantized linear models and extend the framework to nonlinear settings—incorporating non-convex optimization and SGD theory—to model proxy-gradient dynamics for both weights and inputs. Results: Our analysis reveals that quantization range primarily governs the duration of the generalization error plateau, while bit-width controls convergence speed. We observe and explain the characteristic plateau-and-drop behavior in generalization error evolution. This work establishes the first analytically tractable, high-dimensional dynamical systems framework for efficient quantized training.
To address the fundamental trade-off between biased gradients induced by the Straight-Through Estimator (STE) and the high computational cost of unbiased zeroth-order (ZO) optimization in quantized neural network training, this paper proposes First-Order-Guided Zeroth-Order Gradient Descent (FOG-ZO). FOG-ZO is the first method to synergistically integrate the directional information from STE with unbiased ZO gradient estimates, enabling low-cost correction of STE bias and substantially reducing reliance on expensive ZO queries. By preserving near-first-order optimization efficiency while improving gradient estimation fidelity, FOG-ZO facilitates efficient quantization-aware pretraining. Experiments on DeiT, ResNet, and LLaMA demonstrate that FOG-ZO achieves up to 8% higher accuracy or a 22-point reduction in perplexity over baselines, with computational overhead reduced by 796× compared to n-SPSA.
本文通过统计学习理论研究了二值激活两层神经网络中直通估计器的稳定性和泛化能力,利用算法稳定性解释其统计泛化,并给出了明确的模型稳定性和泛化界。
Backpropagation relies on two-pass computation and storage of intermediate activations, severely limiting training efficiency for large-scale models; while forward-mode automatic differentiation avoids these bottlenecks, existing gradient estimation methods suffer from high variance and substantial bias, hindering scalability. This paper proposes a novel backpropagation-free gradient estimation framework: it constructs low-bias guess directions by controllably modulating upstream Jacobian matrices and incorporates a bias–variance trade-off analysis to achieve efficient gradient direction approximation in the forward pass. Theoretically, we establish that the intrinsic low-dimensional structure of neural network gradients critically governs estimation quality. Empirically, the method exhibits improved performance with increasing network width, demonstrating superior scalability. It enables lightweight training of ultra-large models, offering a new paradigm for scalable deep learning.
This work proposes a selective gradient computation framework to reduce per-iteration computational costs in deep learning training. By excluding low-loss samples from backpropagation while employing defensive mix sampling and Horvitz–Thompson inverse probability weighting, the method yields unbiased (or controllably biased) gradient estimates. It provides the first theoretical guarantees for loss-based selective backpropagation, elucidating the failure mechanisms of uncompensated approaches under noisy or imbalanced data and establishing convergence and bias bounds. Experiments on multiple real-world datasets demonstrate 28%–54% reductions in gradient computations with no statistically significant degradation in test performance compared to full-batch SGD (p ≥ 0.5). Under extreme class imbalance, the proposed method achieves an AUC of 0.9991, substantially outperforming uncompensated baselines, which attain only 0.53–0.62 AUC.
Binary neural networks (BNNs) suffer from reliance on floating-point arithmetic during backpropagation, hindering fully binary training. To address this, we propose Binary Error Propagation (BEP), the first method to formulate a discrete-chain rule for error propagation using purely binary vectors to represent gradient signals. BEP performs both forward and backward passes exclusively via bitwise operations—eliminating all floating-point computations and full-precision parameters. This enables end-to-end fully binary training, with particular efficacy for recurrent neural network (RNN) architectures. Experiments demonstrate significant accuracy improvements: +6.89% on multilayer perceptrons and +10.57% on RNNs, respectively, over baseline BNNs. The implementation is publicly available.
This work proposes a randomized, unbiased vector-Jacobian product (VJP) approximation method with minimal variance to replace exact computations in backpropagation, aiming to reduce the computational and memory costs of training deep neural networks. The approach achieves theoretically optimal estimation under sparsity constraints and establishes a principled trade-off between approximation accuracy and per-iteration training cost. Empirical evaluations on multilayer perceptrons, BagNets, and Vision Transformers demonstrate that the method substantially lowers training overhead while preserving model accuracy almost entirely.
This work addresses the challenges of vanishing and exploding gradients in backpropagation by proposing a gradient-free training method for deep neural networks. Built upon an extremely simple Monte Carlo strategy, the approach randomly perturbs network parameters and retains updates that reduce the loss, enabling effective optimization on a single GPU. The study demonstrates, for the first time, that such a minimalist gradient-free algorithm can directly train networks exceeding 20 layers without relying on batch normalization or residual connections. Moreover, it is compatible with purely pruned architectures, discrete weights, and non-standard activation functions such as Gaussian activations. Experiments validate the method’s efficacy and generality by successfully training ultra-deep networks, wide networks with 16,384 neurons, and a minimal Transformer on MNIST and Tiny Shakespeare benchmarks.
This work addresses the gradient mismatch and information loss in binary neural networks caused by non-differentiable binarization operations. The authors propose a learnable gradient compensation framework that decouples forward and backward gradient flows through a Dual-Path Gradient Compensator (DPGC) and dynamically balances branch contributions via an Adaptive Gradient Scaler (AGS). By integrating auxiliary backpropagation with output decomposition, the method enables more accurate gradient estimation. Notably, this is the first approach to incorporate a theoretically grounded, learnable gradient adaptation mechanism into binary network training. Extensive experiments demonstrate consistent and significant improvements over state-of-the-art methods across diverse tasks, including image classification, object detection, and language understanding.