Score
Designs and analyzes gradient-free optimization algorithms that improve parameters, policies, or other black-box decision variables using only objective (function) evaluations rather than gradients, e.g., via random or finite-difference parameter perturbations and performance estimation from sampled evaluations. Focuses on sample-efficient, perturbation-based update rules and trust-region–style adjustments to accelerate convergence with minimal evaluations.
This paper investigates the convergence of gradient-free optimization algorithms under noisy, nonsmooth, or both settings—typical in black-box optimization where gradients are unavailable. We propose a unified analytical framework integrating model-based strategies and smoothing techniques, replacing gradient estimation for nonsmooth or stochastic objectives with smooth approximations via a generalized gradient descent recursion. Under minimal regularity assumptions—requiring only local bounded variation—we rigorously establish convergence guarantees for both deterministic and stochastic settings. Our analysis uncovers a fundamental trade-off between the smoothing parameter and step size, characterizing their joint impact on convergence rate and stability. Extensive experiments on diverse machine learning classification tasks demonstrate the method’s effectiveness and robustness against noise and nonsmoothness.
This paper addresses derivative-free optimization (DFO) under noisy conditions, systematically investigating the fundamental trade-off between sample efficiency and gradient estimation accuracy in finite-difference (FD) methods versus Kiefer–Wolfowitz (KW) and simultaneous perturbation stochastic approximation (SPSA) algorithms. Through extensive empirical evaluation, we demonstrate—for the first time—that high-accuracy, batched FD gradient estimators combined with standard gradient descent consistently outperform classical KW and SPSA across low- to high-dimensional noisy optimization tasks, achieving both faster convergence and higher solution accuracy. Our key contribution is the empirical validation of the batched FD framework as a superior paradigm for noisy DFO, grounded in its favorable variance–sample-size trade-off: batched FD attains lower gradient estimation variance per sample than stochastic approximation methods, enabling more reliable descent directions and improved overall optimization performance.
Existing parameter-free stochastic optimization methods still rely on prior bounds of problem-specific parameters, hindering truly assumption-free optimization. This work proposes Grasp, a universal framework that integrates self-bounding analysis to automatically determine the parameter search range, thereby achieving the first fully prior-knowledge-free stochastic optimization algorithm. The method enjoys near-optimal convergence guarantees in both non-convex and convex settings: in the non-convex case, it attains the optimal convergence rate up to logarithmic factors, while in the convex case, it simultaneously achieves acceleration and universality. Furthermore, by modeling interpolation variance, the approach provides novel theoretical guarantees for model ensembling, matching the performance of existing methods that require meticulous hyperparameter tuning.
This paper addresses Lipschitz continuous, nonsmooth, nonconvex stochastic optimization in decentralized networks—without requiring gradient information. We propose two zeroth-order distributed algorithms: DGFM and its enhanced variant DGFM+. DGFM+ is the first method to integrate randomized smoothing, gradient tracking, and variance reduction in a decentralized zeroth-order setting, incorporating a novel double-batch sampling scheme that improves the convergence complexity to $O(d^{3/2}delta^{-1}varepsilon^{-3})$. Theoretically, both algorithms are proven to converge to an $(delta,varepsilon)$-Goldstein stationary point. The framework supports flexible oracle queries—including single-sample, mini-batch, and periodic large-batch evaluations. Empirical results on real-world datasets demonstrate that DGFM+ significantly outperforms existing decentralized zeroth-order methods in terms of both convergence speed and solution quality.
Learned optimizers (L2Os) suffer from poor out-of-distribution generalization, limiting their applicability beyond the training data distribution. Method: This paper proposes a novel paradigm integrating classical optimization priors with data-driven modeling. It systematically incorporates fundamental optimization principles—specifically scale invariance and affine covariance—into the architecture design. We introduce a parameterized quasi-Newton update module explicitly constrained to preserve BFGS structure, and jointly optimize it via end-to-end training that unifies optimization-theoretic modeling, neural network architecture design, and meta-learning. Contribution/Results: The resulting enhanced BFGS algorithm significantly outperforms both standard L2Os and conventional solvers on unseen problem classes, dimensions, and condition numbers. It achieves over 40% improvement in cross-distribution generalization performance, establishing a new pathway toward more transferable and robust learned optimizers.
This work addresses high-dimensional black-box global optimization under noisy function evaluations and unknown local smoothness. It proposes and implements the Parallel Optimistic Optimization (POO) algorithm, which dispenses with prior knowledge of the target function’s smoothness near the optimum. By integrating multi-scale search with an adaptive mechanism through a parallel optimistic optimization strategy, POO achieves robust and efficient optimization without requiring smoothness assumptions. Theoretical analysis shows that after $n$ function evaluations, POO incurs an optimization error at most a $\sqrt{\ln n}$ factor worse than that of the best-known algorithm assuming known smoothness. This result eliminates the traditional reliance on prior smoothness information, thereby extending applicability to a broader class of challenging optimization problems.
This work addresses the lack of theoretical guarantees for Schedule-Free optimization methods in non-convex settings, particularly regarding convergence and saddle-point escape. By constructing a continuous-time limit that corresponds to a non-autonomous ordinary differential equation, the authors develop a Lyapunov-based analytical framework. Within this framework, they establish—for the first time—that the standard Schedule-Free gradient descent and its stochastic variant achieve the optimal worst-case convergence rate among first-order methods, without requiring algorithmic modifications or strong assumptions. Moreover, they rigorously prove that these methods avoid strict saddle points. Bridging non-convex optimization, ODE modeling, and non-autonomous dynamical systems theory, this study provides the first comprehensive theoretical foundation for both convergence and saddle-point escape in Schedule-Free methods.
This work addresses the poor sample efficiency of Monte Carlo estimators in derivative-free stochastic nonconvex optimization by proposing a hierarchical adaptive sampling strategy within a trust-region framework. The method dynamically controls the Monte Carlo estimation error, ensuring it remains below a stationarity threshold derived from the trust-region radius, thereby achieving high accuracy while substantially reducing sample complexity. Theoretical analysis establishes an improved sample complexity bound for the proposed algorithm, and numerical experiments further demonstrate its computational efficiency and optimization accuracy in high-dimensional settings.
This study addresses the limited statistical gains achievable by existing data-driven optimization methods in the absence of informative priors. It provides a systematic analysis of the “directional perturbation” empirical optimization (EO+) framework, establishing for the first time a unified theoretical perspective that demonstrates only second-order improvements are attainable without geometrically valid side information—thereby confirming a “no free lunch” negative result. The work breaks new ground by showing that incorporating geometrically effective side information and substantially increasing a key hyperparameter enables first-order statistical improvement. Furthermore, it introduces a gain-maximization mechanism based on excess risk estimation, significantly enhancing decision efficiency. By integrating bootstrap resampling, control variates, distributionally robust optimization, and transfer learning, this research bridges data-driven optimization with Monte Carlo variance reduction theory.
This work proposes PF-AGD, a deterministic accelerated first-order algorithm for smooth nonconvex optimization that operates without prior knowledge of the smoothness constant. By integrating adaptive backtracking with a gradient-driven restart mechanism, PF-AGD dynamically estimates local curvature on the fly, thereby eliminating reliance on theoretical parameter tuning. Notably, it achieves the best-known first-order oracle complexity of $\tilde{O}(\varepsilon^{-5/3})$ without requiring any prespecified smoothness constant—a first in the literature. Empirical evaluations demonstrate that PF-AGD outperforms existing parameter-free methods, including practical variants of AGD-Until-Guilty, and matches the performance of nonlinear conjugate gradient methods, thus offering both theoretical optimality and practical efficacy.