Score
Design, build, and analyze optimization methods that use only ordinal or pairwise comparison queries (e.g., preferences or winner/loser outcomes) instead of numeric function values to locate optima or epsilon-stationary points and to navigate nonconvex objective landscapes. Prove convergence guarantees and derive query-complexity and robustness bounds for algorithms that rely solely on comparison information.
This work studies zero-order stochastic optimization relying solely on ordinal feedback (e.g., pairwise preferences), motivated by practical settings such as human-in-the-loop reinforcement learning. For stochastically smoothed objective functions, we propose the first rank-based zero-order optimization algorithm and establish, for the first time, explicit, non-asymptotic query complexity bounds. Theoretical analysis shows that, under both convex and non-convex assumptions, the algorithm achieves query efficiency matching that of optimal value-based methods; moreover, purely ordinal information suffices to attain information-theoretic optimality under stochastic smoothness. Departing from conventional drift analysis and information-geometric frameworks, we introduce novel analytical tools to characterize the information capacity of ordinal feedback. To our knowledge, this is the first work providing rigorous, non-asymptotic theoretical guarantees for preference-driven optimization.
This paper investigates the computational complexity of finding (δ,ε)-stable points in stochastic nonsmooth nonconvex optimization. To overcome the bottleneck of the prior best-known complexity O(ε⁻⁴δ⁻¹), we establish, for the first time, a rigorous reduction framework from nonsmooth nonconvex optimization to online learning, reformulating the problem as an optimistic online learning task and integrating stochastic subgradient methods with complexity lower-bound analysis. Our contributions are: (1) achieving the optimal stochastic gradient query complexity O(ε⁻³δ⁻¹); (2) deriving a tight lower bound that certifies its theoretical optimality; (3) naturally extending the analysis to the second-order smooth setting, yielding a new complexity bound O(ε⁻¹·⁵δ⁻⁰·⁵); and (4) unifying and recovering all known optimal or state-of-the-art results for smooth and higher-order smooth settings—thereby establishing a paradigm-level unification.
In zeroth-order (ZO) optimization, a fundamental trade-off exists between the number of queries per iteration and the total number of iterations under a fixed query budget; its optimal allocation mechanism has remained unresolved. Method: We systematically model and solve this multi-query allocation problem, proposing two gradient aggregation strategies: ZO-Avg (simple averaging) and ZO-Align (a novel projection-based alignment method). Contribution/Results: Theoretical analysis shows that ZO-Avg’s convergence rate does not improve with increasing per-iteration queries under strong convexity, convexity, non-convexity, or stochastic settings. In contrast, ZO-Align significantly enhances gradient estimation accuracy via subspace alignment, yielding monotonically improved convergence rates as per-iteration queries increase. Empirical results confirm the performance gap and reveal that the so-called “multi-query paradox” stems from suboptimal aggregation—not inherent limitations. This work establishes, for the first time, the theoretical optimality criterion for query allocation under fixed budgets, providing principled guidance for ZO algorithm design.
This work systematically uncovers the decisive role of problem geometry—specifically, the curvature of the constraint set and the structure of gradients—in governing the statistical-computational trade-offs of stochastic and online optimization algorithms. We introduce the first geometric measure quantifying the deviation of a constraint set from quadratic convexity, rigorously identifying the geometric origins of suboptimality in subgradient methods. We prove that diagonal-preconditioned SGD achieves minimax-optimal convergence rates under quadratic convex constraints. For non-Euclidean, non-quadratically-convex domains—such as ℓₚ-balls with p < 2—we establish tight convergence bounds for mirror descent and adaptive gradient methods, and uncover, for the first time, a precise correspondence between their convergence rates and the accuracy-computation trade-off in Gaussian sequence estimation. Our results provide geometric criteria for algorithm selection and unify the understanding of when nonlinear updates—e.g., via mirror descent—are necessary to attain statistical optimality.
This work addresses the computational complexity of finding an ε-approximate stationary point in nonconvex–nonconcave minimax optimization problems defined over the unit hypercube. Under the standard oracle model where only function values and gradients are accessible, the paper establishes the first exponential lower bound on query complexity: the number of required oracle queries grows exponentially either in the inverse accuracy parameter 1/ε or in the problem dimension. This result rigorously demonstrates the intrinsic difficulty of nonconvex–nonconcave minimax optimization, implying that no polynomial-time algorithm can exist for this general setting.
This work addresses the problem of efficiently finding ε-stationary points in non-convex optimization under the comparison oracle model, where only comparisons between function values are accessible. The authors introduce, for the first time, a Hessian-normalized estimation subroutine based solely on comparison queries, leveraging both Lipschitz continuity of the gradient and structural properties of the Hessian. Building upon this subroutine, they design both classical and quantum algorithms: the classical algorithm requires Õ(n²/ε¹·⁵) comparison queries, while the quantum algorithm exploits quantum superposition to reduce the query complexity to Õ(n/ε¹·⁵). This quantum method constitutes the first algorithm capable of locating ε-stationary points in the comparison oracle setting, achieving a significant improvement in query efficiency over classical approaches.
This work addresses the long-standing lack of non-asymptotic theoretical guarantees for ranking-based zeroth-order (ZO) optimization methods—such as CMA-ES and natural evolutionary strategies. We establish, for the first time, explicit non-asymptotic query complexity bounds for top-$k$ direction selection algorithms. Departing from conventional drift analysis and information-geometric approaches, we introduce a novel analytical framework that uncovers the fundamental mechanisms enabling efficient optimization under ranking feedback. For $mu$-strongly convex $L$-smooth functions, our bound is $ ilde{O}ig((dL/mu)log(1/varepsilon)ig)$; for nonconvex $L$-smooth functions, it is $Oig((dL/varepsilon)log(1/varepsilon)ig)$, holding with probability at least $1-delta$. These results provide the first rigorous quantification of the robustness–efficiency trade-off induced by ranking feedback, thereby filling a critical theoretical gap in ranking-based ZO optimization.
This paper studies linear combinatorial optimization in the comparison oracle model, where the optimal solution is identified solely via pairwise comparisons ( <, =, > ) of weights of feasible subsets. To address the overly strong assumptions of traditional value oracles, we propose the first global subspace learning framework, integrating inferential dimension analysis, algebraic structural modeling, and discrete integer sorting techniques. Theoretically, we achieve the first separation between information complexity and computational complexity. Algorithmically, our approach attains an $O(nB log(nB))$ query complexity for fundamental combinatorial structures—including minimum cut, shortest path, and bipartite matching—where $n$ denotes the ground set size and $B$ the bit complexity of weights. This establishes a polynomial-time, low-query paradigm for weakly supervised combinatorial optimization, significantly advancing beyond prior value-oracle–dependent methods.
This work addresses the incompatibility between traditional static Nash equilibria and individual regret in online dynamic games. To bridge this gap, the authors introduce two new performance measures: the Static Duality Gap (SDual-Gap) and the Dynamic Saddle-Point Regret (DSP-Reg). They develop a unified theoretical framework applicable to strongly convex–strongly concave functions, min-max exponentially concave functions, and those satisfying the two-sided Polyak–Łojasiewicz condition. By reducing the problem to classical online convex optimization (OCO), they design corresponding algorithms and establish tight theoretical bounds for the proposed metrics. The framework not only unifies the analysis across diverse function classes but also demonstrates practical efficacy through applications such as two-player portfolio selection, confirming its generality and real-world relevance.