EvE: An Alternate Optimizer to Adam

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive evaluation costs and difficulty of rapid pruning associated with Adam in hyperparameter and architecture search. To overcome these limitations, this work proposes the EvE optimizer as an efficient surrogate. The method integrates steady-state differential evolution with a directed Adam fallback mechanism, incorporating gradient-aware strategies to enable low-cost candidate generation. Furthermore, it provides theoretical guarantees of monotonically non-increasing performance under greedy selection. Experimental results demonstrate that EvE matches or surpasses Adam on most benchmarks while achieving a 1.7× to 3.5× acceleration in search speed, thereby significantly expediting the model selection pipeline.
📝 Abstract
Adam and its variants dominate neural network training, but a single run only reveals whether a configuration works well after most of its budget is spent, a poor fit for hyperparameter or architecture search, where configurations must be ranked cheaply and pruned early. We introduce EvE (Evolutionary Explorer), a steady-state, population-of-four differential evolution (DE) optimizer with a targeted Adam fallback: each iteration proposes one candidate via DE, running a short burst of gradient descent only if the DE step fails to improve on the incumbent. Selection is greedy, so on a deterministic objective the best-so-far value is provably monotone non-increasing, and since gradients are used only as a targeted rescue, per-iteration cost stays within a constant factor of a single Adam step regardless of dimension. Under a fixed, evaluation-cost-matched budget, EvE wins or ties Adam on 76% of 70 (problem, dimension) cells across seven scalable benchmarks up to one million variables. On three real neural-network tasks (an MLP on MNIST, and LoRA fine-tuning of a 1.5B-parameter language model on two datasets) EvE finishes the same charged budget 1.7-3.9x faster, at a modest cost in final quality (about one accuracy point on MNIST, 9-11% higher relative test loss on the two fine-tuning tasks; on GSM8K, Adam is about 5 accuracy points more accurate, and fine-tuning lowers accuracy below the base model for both). Inside successive halving on UCI Adult, EvE completes hyperparameter and architecture searches 3.1-3.5x faster, ranking configurations about as consistently with Adam as Adam does with itself across seeds (Kendall's tau 0.66-0.69). EvE is not a total replacement for Adam as a final-stage trainer, but a fast, gradient-aware proxy for the search-heavy, budget-constrained regime one level up.
Problem

Research questions and friction points this paper is trying to address.

hyperparameter search
architecture search
budget-constrained optimization
neural network training
early pruning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Differential Evolution
Adam Fallback
Hyperparameter Search
Successive Halving
Proxy Optimizer
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.