Score
Designs and implements organized benchmarking competitions: defines tasks and datasets, specifies evaluation protocols and metrics, builds submission pipelines, scoring and leaderboard systems, and provides baseline models and reproducible evaluation artifacts. Operates the challenge by writing rules and governance, validating and scoring submissions, managing participant onboarding and data access, and producing post‑challenge analysis and reports.
Organizing periodic algorithm competitions poses significant challenges, including cumbersome submission management, poor cross-platform compatibility, and non-reproducible evaluations. This paper proposes a scalable, automated online competition system that integrates a web service architecture, task queues, and Docker-based containerization to fully automate submission ingestion, isolated execution, and automatic grading. Its key contribution is a lightweight, container-based environment isolation mechanism that ensures evaluation fairness and result reproducibility while enabling unified assessment across heterogeneous development environments. The system has been successfully deployed in multiple international competitions—including the Grid-Based Pathfinding Competition and the League of Robot Runners—demonstrating substantial reductions in organizational overhead, improved grading efficiency and accuracy, and robust support for longitudinal tracking of algorithmic progress. It establishes a sustainable, production-grade technical infrastructure for competitive algorithm evaluation.
AI challenge outcomes frequently suffer from fragmentation, poor reproducibility, and diminishing scholarly impact post-competition. To address this, we propose a systematic framework for sustaining challenge influence. First, we define target stakeholders and sustainable translation pathways. Second, we design the first standardized “post-challenge paper” template to structure reporting, evaluation results, and dissemination activities. Third, we establish a methodology for transforming challenge outputs into enduring benchmarks or academic resources—integrating knowledge organization, interactive visualization, open science practices, and community-driven outreach. The resulting reusable *Post-Challenge Work Guide* has enabled multiple AI challenges—including MedPerf and BraTS—to evolve into authoritative public benchmarks. This framework significantly improves result reproducibility, cross-institutional collaboration efficiency, and academic citation rates.
This work addresses a critical flaw in current mainstream benchmarks: their incentive structures encourage developers to overfit to leaderboard rankings—a practice known as “benchmaxxing”—thereby obscuring true model capabilities. For the first time, the authors frame benchmark evaluation as a Stackelberg game between the benchmark designer and multiple developers, enabling a formal game-theoretic analysis of how different evaluation protocols shape developer strategies and resulting rankings. Theoretical analysis reveals that existing protocols generally lack a Nash equilibrium, leading to unstable or misleading rankings. In contrast, the proposed “tune-before-test” mechanism is shown to guarantee a unique Nash equilibrium under mild conditions, ensuring that leaderboard rankings faithfully reflect the underlying quality of models.
Existing evaluation of generative AI is hindered by the scarcity of high-quality benchmarks, whose manual construction is costly and time-consuming. Method: We propose the first automated benchmark construction framework powered by collaborative large language model (LLM) agents, decomposing benchmark creation into four sequential stages—planning, generation, verification, and evaluation—integrating task decomposition, agent coordination, human-in-the-loop feedback, and explicit constraint-satisfaction assessment. Contribution/Results: The framework significantly enhances data diversity and metric reliability. Leveraging it, we construct the first high-quality benchmark specifically targeting planning and constraint-satisfaction capabilities in text generation. We systematically evaluate seven state-of-the-art models, uncovering shared failure modes and fine-grained capability disparities. Our work establishes a scalable, reproducible paradigm for evaluating generative AI capabilities, advancing both benchmark methodology and empirical analysis.
Existing benchmarks for knowledge work evaluation largely adhere to traditional NLP task paradigms, failing to capture systems’ capabilities in real-world knowledge-intensive settings. This work proposes a three-step framework—explicitly defining work activities, establishing realistic test environments, and focusing evaluation on deliverable outputs—and derives 18 core knowledge work activities from the O*NET database. Innovatively integrating role responsibilities, local tool usage, and downstream usability into benchmark design, the approach establishes a coherent “work activity–test setup–scoring artifact” alignment. Validation through three case studies (GDPval, OfficeQA Pro, and APEX-SWE) exposes critical misalignments in current benchmarks between tasks, environments, and actual work objectives, offering a new paradigm for evaluating knowledge work systems in practical, application-oriented contexts.
研究通过审计254个SWE-bench提交,发现编码代理排行榜上的微小差异不能准确反映系统优劣,建议报告比较集分辨率和模型-框架来源。
Traditional benchmarks provide only aggregate scores, offering insufficient evidence to support reliable deployment decisions and thereby creating a disconnect between evaluation and action. To address this gap, this work proposes a “deployment-completeness” benchmarking framework, introducing novel metrics—evidence fibers, completeness curves, and certifiable proportions—alongside a systematic audit methodology comprising evidence fiber analysis, response ranking intervals, conformal coverage evaluation, and a certify-then-acquire decision pipeline. Empirical evaluation on benchmarks such as Tox21, Matbench, and JARVIS reveals that conventional approaches suffer a drastic drop in channel coverage to 10.07% under real-world deployment conditions. In contrast, the proposed method reduces error-driven deployment decisions to 0.027% on Tox21 and 0.128% on JARVIS, substantially enhancing deployment reliability.
This study addresses a critical gap in current AI evaluation methodologies, which often overlook the impact of low-resource deployment conditions—such as noisy inputs, limited hardware capabilities, and unstable network connectivity—on system usability. The work proposes a novel evaluation framework that treats the deployed system as the unit of assessment, integrating task performance with real-world deployment contexts across multiple dimensions. Departing from conventional leaderboard-based approaches, the framework tailors evaluation criteria to specific application categories and introduces a standardized reporting system comprising benchmark cards, deployment profiles, and failure-handling mechanisms. By balancing comparability with contextual sensitivity, this approach provides policymakers and practitioners with clear, actionable insights for informed AI deployment decisions.
Current large-model agent benchmarks rely on static average-score leaderboards, which poorly predict real-world performance in out-of-distribution deployment scenarios. This work proposes a deployment-oriented, multidimensional evaluation framework centered on predictive validity—the correlation between in-sample and out-of-distribution rankings. We introduce a novel evaluation paradigm grounded in predictive validity, featuring a twelve-tier measurement architecture and three falsifiable criteria for out-of-distribution assessment. Through a preregistered empirical study integrating 14 parallel implementations, seven existing benchmarks, and extensions across multimodal settings, diverse agent orchestrations, and retrieval-augmented approaches, we demonstrate that conventional leaderboards yield unstable rankings under distribution shift, whereas our paradigm exhibits significantly stronger predictive power, thereby establishing a methodological foundation for next-generation agent benchmarks.
This study addresses a critical gap in existing tool-calling evaluation benchmarks: the lack of validation of the evaluators themselves, which risks conflating assessment artifacts with agents’ true capabilities. Through a systematic audit of four prominent benchmarks—BFCL v4, τ2-Bench, LiveMCPBench, and MCP-Atlas—the authors conduct expert review of 496 tasks, replicate experiments, and perform trajectory-level analysis, revealing an 18.5% disagreement rate between automated evaluators and human judgment. Notably, LiveMCPBench exhibits a score variance of up to 18.9 percentage points upon re-evaluation, sufficient to overturn leaderboard rankings. To address these issues, the work introduces the first unified taxonomy of tool-calling evaluation failures, advocates for distinct measurement of tool invocation, task completion, and result verification, and releases Tool-Veritas—a configurable benchmark—and Harness Lab, an open-source evaluation platform.