🤖 AI Summary
This study addresses the unavailability of gradients in black-box optimization and the inefficiency of conventional zeroth-order methods by proposing POEM-CMA, a parameter-free stochastic zeroth-order optimization algorithm. The method introduces an empirical effective dimension in place of the ambient dimension to concentrate on informative directions, and integrates covariance matrix adaptation with alignment techniques to enable anisotropic sampling, thereby efficiently accommodating problems with low-rank structures. Theoretically, this work establishes that the proposed algorithm achieves a near-optimal convergence rate. Empirically, experimental results on LibSVM datasets demonstrate that POEM-CMA significantly outperforms the original POEM method.
📝 Abstract
Zeroth-order optimization methods are essential for solving black-box problems where gradient information is unavailable or expensive to compute. This paper presents POEM-CMA, a novel parameter-free stochastic zeroth-order algorithm that extends the recent POEM method by integrating covariance matrix alignment and the notion of effective dimension.
In contrast to traditional zeroth-order approaches that rely on isotropic random directions, POEM-CMA performs anisotropic sampling by constructing a covariance matrix from gradient estimates. This enables the algorithm to focus sampling efforts on the most informative directions. We introduce the use of the empirical effective dimension $d^* = \frac{\operatorname{tr}(\hatΣ)}{λ_{\max}(\hatΣ)}$, which reflects the intrinsic dimensionality of the problem and replaces the ambient dimension in both sampling and complexity analysis.
We prove that POEM-CMA achieves a near-optimal convergence rate, requiring only $\tilde{\mathcal{O}}\left(\frac{d^* κ(\hatΣ) L^2 D_{\mathcal{X}}^2}{\varepsilon^2}\right)$ stochastic zeroth-order oracle queries. The method remains fully parameter-free and demonstrates significant improvements over the original POEM in problems with low-rank structure where $d^* \ll d$. Numerical experiments on hinge-loss binary classification tasks using LibSVM datasets confirm the practical superiority of the proposed approach.