🤖 AI Summary
This study addresses the rapid saturation of static benchmarks for large language models and the limitations of existing evaluation frameworks constrained by hard-coded rules. To overcome these challenges, this work proposes a dual-loop search framework in which an inner loop automatically generates evaluation benchmarks while an outer loop leverages historical trajectory analysis to perform end-to-end iterative optimization of the benchmark generation workflow through a meta-orchestration mechanism, thereby transcending the constraints of isolated task evolution. Experimental results demonstrate that the proposed approach significantly enhances benchmark difficulty, discriminative power, and evaluator robustness in domains such as competitive programming. Ultimately, this framework establishes a novel paradigm for dynamic evaluation systems capable of continuously adapting to advancing model capabilities.
📝 Abstract
Rapid progress in Large Language Models (LLMs) is saturating static benchmarks faster than they can be designed. While existing automated evolution frameworks attempt to generate harder questions by perturbing individual tasks, they remain constrained by rigid, hard-coded generation rules. Moving beyond the evolution of isolated tasks, we propose to optimize the benchmark generation workflow itself end to end with MetaBench-Harness, a dual-loop search framework. Specifically, the inner loop utilizes a benchmark harness to generate a new benchmark in each round, while the outer meta-harness orchestration layer iteratively refines and searches over harness implementations based on historical evolution trajectories. By applying MetaBench-Harness to the competitive programming CodeContests and Olympiad mathematics AIME-2024 datasets, we demonstrate that the evolved benchmarks are challenging and discriminative for frontier models. Trajectory and quality analyses verify that MetaBench-Harness enables multi-dimensional evolution, steadily improving evolution reasonableness, benchmark competency, and evaluator robustness across successive rounds. Furthermore, case studies reveal its effective utilization of diverse difficulty levers to reframe problems and elevate required capabilities. Ultimately, this work provides a solution to the pressing challenge of benchmark saturation.