Score
Designs and analyzes optimization algorithms and frameworks that intentionally limit the number of adaptive rounds—often by grouping queries into batches—to perform optimization with few sequential interactions. These methods aim to trade adaptivity for parallelism or efficiency, reducing sample complexity and running time while providing provable approximation or performance guarantees.
This work addresses the weak theoretical convergence guarantees of adaptive optimization algorithms—particularly those employing moving-average momentum estimators—in nonconvex optimization. Methodologically, it establishes a unified convergence analysis framework: (i) it provides the first rigorous proof that monotonic increase of the first-order momentum parameter ensures convergence; (ii) it uncovers a phased co-adaptation mechanism between momentum and step size that enables acceleration; and (iii) it extends the analysis to composite, minimax, and bilevel optimization settings. Theoretical contributions include: (i) nonconvex convergence guarantees applicable to broad classes of adaptive methods (e.g., Adam-type); and (ii) novel, efficient minimax and bilevel optimization algorithms that avoid large batch sizes or double-loop schemes. Empirical results confirm improved convergence rates and generalization performance, corroborating the theoretical insights.
Existing optimizer benchmarks are often confined to a single batch size, overlooking the reliability of hyperparameter scaling rules across varying batch sizes. This study systematically investigates the dependence of optimizer performance on batch size through large-scale language model pretraining experiments, comparative evaluations of multiple optimizers, and extensive hyperparameter tuning. It provides the first empirical evidence that no universal Muon scaling rule generalizes across training configurations, and that the optimal optimizer shifts dynamically with batch size. By exposing the failure of prevailing scaling assumptions, this work demonstrates that optimizers must be independently selected and tuned for specific batch sizes to effectively enhance training efficiency.
本文针对自适应数据分析中因查询链导致的泛化误差累积问题,提出一种基于加权依赖图的程序分析方法来近似计算并控制自适应性。
To address the inefficiency and poor generalizability of manual hyperparameter tuning—particularly for learning rates—this paper proposes a dynamic online meta-optimization framework that formulates learning rate adaptation as a discounted cumulative regret minimization problem over time. The method employs a gradient-based meta-update mechanism, enabling plug-and-play integration with any first-order optimizer (e.g., SGD, Adam) to achieve decoupled, real-time, adaptive step-size optimization. Key contributions include: (i) the first formalization of meta-optimization as discounted regret minimization; and (ii) a low-complexity variant that preserves theoretical rigor while ensuring computational efficiency and strong generalization. Experiments across diverse tasks demonstrate faster convergence, enhanced robustness to initialization and task heterogeneity, competitive performance against hand-tuned optimal schedulers, and significantly lower computational overhead compared to conventional hyperparameter search methods.
This paper studies the *adaptive parallel complexity* of finding an ε-stationary point in nonconvex optimization—i.e., the minimal number of sequential rounds required to achieve ε-stationarity, assuming polynomially many oracle queries per round. We establish *tight upper and lower bounds* on this complexity in both high-dimensional and constant-dimensional settings. In high dimensions, we prove that gradient descent, cubic-regularized Newton’s method, and p-th-order adaptive regularization methods all achieve the optimal round complexity Ω(ε^{−(p+1)/p}). In constant dimensions, we propose a novel algorithm combining grid search with gradient-flow trapping, attaining O(1) rounds. We further derive a per-round query lower bound of Ω̃(ε^{−(d−1)/2}), confirming its adaptive optimality. Key technical tools include chain-based hardness potential analysis, Lipschitz p-th-order derivative constructions, gradient-flow trapping phenomena, and information-theoretic lower-bound derivation.
This work addresses the performance limitations of single algorithms in black-box optimization due to the absence of prior knowledge. It proposes a sequential portfolio strategy that dynamically allocates computational budget across multiple algorithms, leveraging their complementary strengths and the variance reduction effect observed on individual objective functions to enhance overall performance. Notably, the approach requires no parallelization—operating solely through sequential algorithm invocations—and consistently outperforms single-algorithm baselines while uncovering new potential in restart mechanisms and warm-start strategies. Extensive evaluation on the COCO platform using the BBOB benchmark suite across more than 200 algorithm combinations demonstrates that the proposed method reliably surpasses existing baselines, achieving an average relative performance improvement exceeding 14%.
This work addresses optimization problems with convex constraints whose intersection is difficult to project onto, covering both strongly convex smooth and general nonsmooth convex settings. The authors propose a novel algorithm that integrates stochastic feasibility methods with (sub)gradient descent, wherein each iteration randomly samples a subset of constraints and employs an adaptive Polyak stepsize that requires no prior knowledge of problem parameters, complemented by iterate averaging. Theoretical analysis establishes linear convergence under strong convexity and a worst-case rate of $O(1/\sqrt{T})$ for general convex objectives, while the infeasibility measure decays geometrically almost surely. Numerical experiments on QCQP and SVM tasks demonstrate superior computational efficiency over existing methods, and under specific sampling strategies, the algorithm achieves optimal convergence rates.
This work investigates the feasibility of approximating the size of a maximum matching in a graph under the non-adaptive query model. By leveraging non-adaptive adjacency-list queries, probabilistic analysis, and complexity lower-bound techniques, it establishes—for the first time—that any algorithm achieving an approximation better than \(n^{1/3 - \gamma}\) requires \(\Omega(n^{1+\varepsilon})\) queries, thereby demonstrating the essential role of adaptivity in this problem. Complementing this hardness result, the paper also presents an \(n^{1/2}\)-approximation algorithm that uses only \(O(n \log^2 n)\) queries. These results hold both in the standard non-adaptive query model and in the fixed query tree model, highlighting a sharp separation between adaptive and non-adaptive approaches for graph matching approximation.
This work addresses the challenge of efficiently solving numerous related linear programming (LP) subproblems arising in mixed-integer programming (MIP), particularly in contexts such as strong branching and bound tightening, where traditional methods fail to exploit GPU parallelism effectively. The authors propose a GPU-oriented batched first-order optimization method, reformulating the primal–dual hybrid gradient algorithm into matrix–matrix operations to significantly enhance parallel efficiency. This study represents the first systematic application of batched first-order methods to LP subproblems within MIP solvers, advocating that GPUs should perform core computational tasks rather than merely assist CPU-based heuristics. The approach promotes deeper co-design between MIP algorithms and GPU architectures. Experimental results demonstrate that the proposed method outperforms conventional simplex solvers under specific problem scales and hardware configurations.
This work addresses the problem of efficiently selecting the maximum among $n$ unknown values, each observed only through a single unbiased estimate. The authors propose an adaptive weighted averaging strategy grounded in statistical decision theory, which integrates online learning with batch optimization techniques. The method achieves a “no-regret” guarantee: it is admissible, meaning its worst-case performance never falls below that of uniform random selection, while substantially outperforming baseline approaches when favorable structural assumptions hold. By bridging online and batch learning paradigms, this approach establishes new theoretical bounds for online-to-batch conversion, offering both robustness in adversarial settings and enhanced empirical performance under benign conditions.