Score
Design and build scalable, reproducible benchmarking frameworks and experimental protocols that evaluate and compare optimization algorithms across domains, model types, and training regimes; implement metrics, measurement tooling, and analysis pipelines to quantify performance, compute and memory trade‑offs, and failure modes, and produce actionable selection guidance for optimizer choice under resource and accuracy constraints.
This work addresses the lack of systematic guidance in optimizer selection for large-scale model training, where existing approaches are fragmented and struggle to accommodate diverse computational, memory, and task requirements. The authors propose a five-stage meta-pipeline that unifies the update mechanisms of optimizers and introduce a norm-constrained Linear Minimization Oracle (LMO) to achieve geometric unification. Building on this foundation, they establish a novel two-dimensional taxonomy that jointly considers optimizer families and optimization objectives, along with a comprehensive benchmarking framework spanning multiple domains, model scales, and tasks. Systematic evaluation reveals critical trade-offs among optimizers in terms of performance, robustness, and efficiency, providing a practical coordinate system and empirical basis for both optimizer selection and future algorithm design.
Existing black-box optimization benchmarking frameworks suffer from dependency conflicts, tight environment coupling, and poor integration scalability, severely compromising experimental reproducibility. To address these challenges, we propose Bencher—a modular benchmarking framework introducing the novel concept of *benchmark isolation*: it decouples benchmark execution from optimization logic via lightweight, version-agnostic RPC abstractions and isolated virtual environments. Bencher supports cross-platform deployment across local machines, Docker containers, and HPC systems (via Singularity). It unifies modeling of heterogeneous search spaces—including continuous, categorical, and binary domains—enabling reproducible benchmark integration across environments, software versions, and application domains. The framework currently supports one-click integration of 80 real-world benchmarks, with a zero-configuration, lightweight client. Empirical evaluation demonstrates substantial improvements in reliability, reproducibility, and engineering compatibility for black-box optimization algorithm assessment.
This paper identifies systemic flaws in benchmark evaluation for algorithm selection—particularly in black-box optimization. First, leave-instance-out cross-validation induces data leakage, yielding spuriously high accuracy. Second, scale-sensitive performance metrics (e.g., raw function-value error) introduce substantial optimistic bias. Through empirical analysis, ablation studies, and meta-model diagnostics, the authors quantitatively demonstrate that non-informative features can achieve >90% accuracy under flawed evaluation protocols, and unnormalized objective functions overestimate meta-model performance by over 40%. To address these issues, they propose a decoupled evaluation framework that separately assesses *feature effectiveness* and *metric-induced interference*, emphasizing objective-function normalization and error attribution analysis. This work advances algorithm selection evaluation toward standardization, robustness, and reproducibility.
Existing mainstream benchmark suites (e.g., BBOB, CEC) lack fidelity to real-world continuous and mixed-integer optimization problems—failing to reflect their structural characteristics, practical constraints, and information limitations—leading to misuse in algorithm competitions, automated algorithm selection, and industrial decision-making. Method: We propose a next-generation, scenario-driven continuous optimization benchmarking framework featuring: (1) a curated benchmark suite grounded in real-world problems; (2) an interpretable high-dimensional problem feature space coupled with an open-source performance database; (3) native support for multi-objective optimization, noisy environments, and algorithm behavioral analysis; and (4) community-driven, dynamic evolution via a collaborative platform. Contribution/Results: The framework bridges the gap between academic evaluation and industrial requirements, significantly enhancing the practicality, interpretability, and reliability of benchmarks for algorithm selection and deployment decisions. It fosters a sustainable, scientifically rigorous, and engineering-ready benchmarking ecosystem.
Fair evaluation of deep learning training algorithms faces three key challenges: inconsistent termination criteria, high workload sensitivity, and difficulty isolating hyperparameter tuning. This paper introduces AlgoPerf—the first time-oriented, multi-workload training algorithm benchmark—featuring robustness-aware workload variant design and a standardized termination protocol, with hyperparameter tuning rigorously isolated. Evaluated on a unified hardware platform, AlgoPerf employs a diverse multi-task workload suite and a systematic optimizer comparison methodology to enable latency-accuracy co-evaluation across models, datasets, and hardware. Experiments reveal substantial latency disparities among mainstream optimizers, establish reproducible state-of-the-art baselines, and deliver the first quantitative, fair, and engineering-practical evaluation standard for training algorithm improvement.
Conventional benchmark suites employ fixed, inflexible functions, limiting systematic and goal-oriented evaluation of optimization algorithms. Method: This paper proposes a通用可配置的 continuous numerical optimization benchmark generator, introducing a novel paradigm based on a single parameterized basis function—replacing traditional multi-function composition. Through parameterized modeling, high-dimensional nonlinear transformations, structured perturbations, and configurable feature mapping, the generator enables full-dimensional, fine-grained, and reproducible control over critical morphological properties—including modality, symmetry, separability, condition number, and basin geometry. Contribution/Results: Generated instances comprehensively span the spectrum of unimodal/strongly multimodal, low-/high-dimensional, and well-/ill-conditioned scenarios. The framework significantly enhances the systematicity, controllability, and interpretability of algorithmic evaluation in continuous optimization.
Existing optimization modeling benchmarks are confined to purely textual inputs, rendering them inadequate for real-world decision-making scenarios that often involve multimodal (text-and-image) information. This work proposes and constructs MOptBench, the first solver-grounded multimodal optimization modeling benchmark, encompassing six problem categories, 26 subcategories, and three difficulty levels. The benchmark ensures correctness through structured instance generation and rigorous validation via exact solvers, enabling fine-grained evaluation and error attribution. Evaluation of nine multimodal large language models on 780 verified instances reveals that the best-performing model achieves a pass@1 rate of 52.1%, while general-purpose models succeed on only 15.9% of hard instances, and math-specialized models fail entirely—highlighting significant limitations in current models’ capacity for complex multimodal reasoning.
Existing benchmarks struggle to evaluate large language models’ (LLMs’) ability to design efficient algorithms for real-world, large-scale optimization problems, often being confined to small-scale or simplified settings. This work proposes FrontierOR—the first large-scale optimization benchmark constructed from top-tier operations research journals—comprising 180 structurally diverse and realistically scaled tasks, accompanied by standardized instances and an expert-validated hidden test set. We assess the algorithm-generation capabilities of seven state-of-the-art open-source LLMs under both single-shot generation and test-time evolution settings. Results show that even the strongest single-shot model outperforms Gurobi on only 31% of tasks; furthermore, despite leveraging a powerful code-centric agent with test-time evolution, success rates on challenging tasks remain at just 50%, highlighting both the significant challenges and untapped potential of LLMs in scalable optimization algorithm design.
Existing benchmarks struggle to evaluate the ability of large language models to perform end-to-end optimization tasks in real-world business settings. This work proposes the first comprehensive end-to-end evaluation benchmark that spans business requirement interpretation, mathematical modeling, algorithm selection, code implementation, and report generation. It introduces three key innovations: business-semantic anti-template traps, cross-module consistency checks, and a dual-layer ORAC validity verification framework, covering core optimization paradigms such as integer programming, robust optimization, stochastic programming, and non-convex optimization. Experiments reveal systematic deficiencies in current models—including omitted constraints and inconsistencies between formulated models and generated code—that remain undetected under conventional single-metric evaluations, thereby demonstrating the necessity and effectiveness of this benchmark for assessing complex, multi-stage optimization workflows.
This work addresses the lack of systematic methodologies in model optimization, which often relies on heuristic choices and struggles to accommodate diverse deployment constraints. It formalizes model compression and acceleration as a constraint-aware multi-objective engineering decision problem, establishing a unified and actionable framework grounded in five key dimensions: data availability, latency, memory footprint, accuracy tolerance, and retraining budget. By integrating techniques such as quantization, pruning, knowledge distillation, parameter-efficient fine-tuning (PEFT), and inference optimization, the study proposes tailored optimization pipelines for four representative industrial scenarios, delivering a reproducible and quantifiable guide for technology selection.
Existing algorithm selection models exhibit limited generalization capabilities in real-world optimization scenarios, struggling to maintain consistent performance across diverse domains. This work presents the first systematic evaluation of cross-domain generalization between synthetic benchmarks (BBOB, CEC) and practical applications—specifically robotic trajectory optimization and UAV path planning—using an algorithm selection framework grounded in problem features and historical performance data, complemented by a carefully designed cross-benchmark experimental protocol. The study uncovers the failure mechanisms and success boundaries of current approaches when deployed in realistic settings, thereby providing crucial empirical insights for developing more robust and universally applicable algorithm selection systems.