Score
Designs and implements evaluation leaderboard systems and the supporting infrastructure that collect and validate participant submissions, run standardized evaluation pipelines, compute and store metrics and baselines, and generate ranked leaderboards. This work includes building reproducible scoring pipelines, dataset and metric versioning, submission APIs and UIs, monitoring, scaling, and access controls to ensure fair, auditable, and maintainable benchmarking.
Manually constructing and maintaining machine learning leaderboards is costly and suffers from inconsistent evaluation standards; existing automatic leaderboard generation (ALG) research lacks a unified problem formulation, hindering cross-study comparison and reproducibility. Method: We propose the first unified conceptual framework for ALG, defining its core task as precise extraction of structured experimental entries—including models, datasets, metrics, numerical results, and contextual metadata—from research papers, coupled with cross-paper alignment for consistency. Our approach integrates literature analysis and task abstraction to design a novel paradigm covering all result types and fine-grained metadata, and introduces a standardized evaluation protocol assessing three dimensions: information extraction accuracy, structural completeness, and cross-paper comparability. Contribution/Results: The work establishes a community-agreed benchmarking guideline, providing both theoretical foundations and practical pathways for automated scientific infrastructure in ML research.
Current leaderboards for large language models rely on static benchmarks that fail to capture the diversity of user needs, and their single aggregate scores obscure performance variations across different prompt types. This work addresses these limitations by analyzing data from the LMArena benchmark, uncovering issues such as thematic skew and ambiguous scoring. To overcome these challenges, the authors propose a user-centered, interactive evaluation paradigm that integrates data slicing, preference modeling, and visual design to enable users to define custom prompt slices and weighting schemes, thereby dynamically exploring model rankings. Qualitative studies demonstrate that this approach significantly enhances evaluation transparency and contextual adaptability, empowering users to select models that best align with their specific requirements.
Organizing periodic algorithm competitions poses significant challenges, including cumbersome submission management, poor cross-platform compatibility, and non-reproducible evaluations. This paper proposes a scalable, automated online competition system that integrates a web service architecture, task queues, and Docker-based containerization to fully automate submission ingestion, isolated execution, and automatic grading. Its key contribution is a lightweight, container-based environment isolation mechanism that ensures evaluation fairness and result reproducibility while enabling unified assessment across heterogeneous development environments. The system has been successfully deployed in multiple international competitions—including the Grid-Based Pathfinding Competition and the League of Robot Runners—demonstrating substantial reductions in organizational overhead, improved grading efficiency and accuracy, and robust support for longitudinal tracking of algorithmic progress. It establishes a sustainable, production-grade technical infrastructure for competitive algorithm evaluation.
Current transferability estimation benchmarks suffer from fundamental flaws—namely, unrealistic fixed model spaces and static performance hierarchies—which severely distort evaluation outcomes; simple, dataset-agnostic heuristics frequently outperform sophisticated metrics, exposing a critical mismatch between benchmark protocols and real-world model selection scenarios. Method: The authors conduct a systematic empirical re-evaluation of mainstream transferability metrics across diverse, realistic model spaces and dynamically varying performance rankings. Contribution/Results: They quantitatively identify and characterize the primary sources of benchmark bias for the first time. Crucially, they demonstrate that the prevailing evaluation paradigm is unreliable, propose a novel benchmarking framework that is realistic, dynamic, and task-aware, and provide both theoretical foundations and practical guidelines for designing robust transferability assessment systems.
Current machine learning leaderboards predominantly rely on manually curated entries, and automated dataset collection efforts typically capture only the best-reported results from papers, lacking comprehensive experimental records and fine-grained metadata. This limits transparent and comparable model evaluation. To address this gap, this work proposes MetaLead—a meticulously annotated, structured dataset that systematically compiles all experimental results reported in research papers, explicitly labeling each experiment’s type (e.g., baseline, proposed method, or its variants) and clearly indicating the separation between training and test datasets. By preserving full experimental context, MetaLead substantially enhances leaderboard transparency and analytical depth, offering a high-fidelity, context-rich benchmark resource for reliable cross-study and cross-domain model performance comparison.
Community-driven scientific workflow ecosystems often struggle to sustain themselves due to ambiguous maintenance and user support mechanisms, particularly in cross-platform collaboration and heterogeneous execution environments. This study presents the first cross-platform empirical analysis of the nf-core ecosystem, systematically examining 15,760 GitHub issues, 35,411 pull requests, and 895 forum discussions. By integrating metadata and textual features into predictive models, the research uncovers significant disparities in maintenance and support activities across platforms and highlights weak explicit linkages among them. The findings reveal that issues, pull requests, and forum posts predominantly serve distinct roles—coordinating maintenance, facilitating code integration, and providing user support, respectively. Moreover, issue actionability, diagnostic evidence, and depth of interaction emerge as critical determinants of resolution efficiency.
This study addresses the lack of rigorous evidence in data contamination controversies surrounding agent leaderboards caused by public benchmarks. We propose a measurement-first auditing framework that introduces a novel three-channel taxonomy—training, retrieval, and pipeline—combined with blinded matched controls, paired bootstrap statistics, and consensus scorer validation for systematic evaluation. Our findings reveal that although leakage events occur across most configurations, they fail to form closed channels; critical benchmark gaps are statistically insignificant, and scorers do not pass validation. By establishing a "fail-closed" principle to standardize contamination evidence criteria, this work demonstrates that existing contamination allegations remain insufficiently substantiated, thereby providing a more rigorous auditing paradigm for agent evaluation.
This study addresses the fragmentation of evaluation criteria for automated research systems and the difficulty of direct cross-task comparison. Employing a systematic literature review, it comprehensively examines evaluation designs across six task categories, including literature synthesis and ideation. By comparing benchmark construction and scoring protocols, this work proposes a complementary evaluation framework encompassing output-level, process-level, and human-subject assessments. It reveals the capability differences reflected by distinct designs and underscores the critical role of calibration specificity and resource budgets in performance interpretation. Furthermore, the project identifies gaps in diagnostic evaluation and provides recommendations for standardized reporting and auditing. Ultimately, these contributions offer practical guidance for benchmark selection and future research design in evaluating automated scientific discovery systems.
This study addresses the incomparability and unfairness in long-horizon agent leaderboards caused by imbalanced inputs, evidence, or budgets. It proposes a combinatorial controllability framework that employs comparison windows to identify sources of discrepancy and establishes admissibility tests to reject invalid comparisons. By leveraging out-of-window distractors to derive theoretical bounds on score differences, the framework enables fairness pre-screening. The evaluation pipeline is further optimized through the BioLitBench benchmark, statistical calibration, and reinforcement learning with stage-wise rewards. Empirically, the approach successfully rejects 11 of 21 unfair comparisons and corrects seven erroneous conclusions. Notably, the proposed SCRIBE model achieves a certified ranking interval of [1, 2] on Qwen3.8, significantly outperforming existing pipelines.
This work addresses the challenge of incomparable, non-reusable, and fragmented AI evaluation results stemming from heterogeneous formats, disparate sources, and inconsistent frameworks. To overcome this, the authors propose the first community-governed unified standard that defines JSON Schemas for both metadata and instance-level evaluation results, alongside a source-agnostic architecture and automated conversion tools supporting 31 diverse evaluation formats. Leveraging this standard, they have constructed a large-scale, standardized database encompassing 22,235 models and 2,273 benchmarks, hosted on Hugging Face with support for crowdsourced contributions. This infrastructure significantly enhances the comparability and reusability of evaluation results while fostering efficient cross-community collaboration.