Score
Designs, implements, and reproduces reference baseline models and evaluation protocols to establish and measure reference performance points; this includes constructing and adapting baseline implementations, running control experiments, and implementing normalization or reimplementation steps needed for fair comparison. Also analyzes and compares baseline performance (including human baselines), measures performance across conditions such as labeled ratios or splits, and documents replication details to support valid baseline comparisons.
Current human baselines in large language model evaluation lack methodological rigor and transparency, undermining the validity of claims such as “superhuman performance.” Method: This paper pioneers the systematic integration of classical measurement theory into AI evaluation, establishing a comprehensive theoretical framework spanning human baseline design, execution, and reporting. It introduces an actionable quality assessment system and a standardized checklist, derived via meta-review–driven framework development, structured checklist design, and empirical systematic auditing. Contribution/Results: Applying this framework to diagnose 115 human baseline studies, we identify pervasive methodological flaws. The resulting open-source audit tool significantly enhances reproducibility, comparability, and accountability in AI evaluation. By grounding benchmarking practice in psychometric principles, our work provides a rigorous methodological foundation for scientifically credible model capability assessment.
Experimental reproducibility in Empirical Software Engineering (ESE) is hindered by a fundamental disconnect between idealized methodological assumptions—e.g., standardized protocols and controlled conditions—and researchers’ actual experimental practices. Method: We conducted a two-year ethnographic study involving participant observation, in-depth interviews, and content analysis of experimental artifacts across diverse ESE research teams. Contribution/Results: We identify four critical dimensions—activity diversity, role distribution, conceptual granularity, and domain perspective—in which real-world experimentation systematically deviates from textbook models. Based on these findings, we propose the first high-fidelity conceptual and process model grounded in empirical research practice, explicitly capturing the “practice gap” underlying irreproducibility. This model provides foundational evidence and design principles for developing next-generation reproducibility-support tools, methodological guidelines, and evaluation frameworks in ESE.
This work addresses the lack of systematic and rigorous performance benchmarking methodologies in programming language research, which has undermined the credibility of evaluation results. To remedy this, the paper introduces a closed-loop methodology—Measure-Explain-Test-Improve—that establishes, for the first time, a structured and reproducible workflow for performance assessment in the field. Integrating systematic experimental design, performance metric analysis, result interpretation, and iterative refinement, the approach emphasizes theoretical grounding and practical rigor at every stage. Its key contribution lies in enabling even researchers with limited empirical experience to conduct reliable and methodologically sound performance evaluations, thereby significantly enhancing the scientific validity and reproducibility of performance analysis in programming language research.
Existing evaluation of generative AI is hindered by the scarcity of high-quality benchmarks, whose manual construction is costly and time-consuming. Method: We propose the first automated benchmark construction framework powered by collaborative large language model (LLM) agents, decomposing benchmark creation into four sequential stages—planning, generation, verification, and evaluation—integrating task decomposition, agent coordination, human-in-the-loop feedback, and explicit constraint-satisfaction assessment. Contribution/Results: The framework significantly enhances data diversity and metric reliability. Leveraging it, we construct the first high-quality benchmark specifically targeting planning and constraint-satisfaction capabilities in text generation. We systematically evaluate seven state-of-the-art models, uncovering shared failure modes and fine-grained capability disparities. Our work establishes a scalable, reproducible paradigm for evaluating generative AI capabilities, advancing both benchmark methodology and empirical analysis.
This study addresses the ambiguity and inconsistency in evaluation criteria for software engineering replication studies, which have led to contradictory interpretations and uncertainty in reported results. Through a systematic review of ten replication studies published between 2021 and 2025, combined with qualitative content analysis, statistical principles, and modeling of measurement uncertainty, this work is the first to uncover the heterogeneity and lack of standardized practices in current evaluation approaches. Building on these insights, the paper proposes a unified evaluation framework that integrates statistical theory, methodological rigor, and measurement theory. Empirical illustration demonstrates that the framework effectively enhances the transparency, consistency, comparability, and reliability of replication studies in software engineering.
This study addresses the growing challenge posed by the widespread involvement of AI agents in software development, which undermines the long-standing assumption that development artifacts are exclusively produced by human professionals—an assumption underpinning traditional software metrics. The work systematically exposes how AI-generated traces compromise the foundational premises of established software measurement practices, thereby threatening the validity of prior empirical conclusions. To confront this issue, the authors propose an AI-augmented, systematic replication methodology that integrates modern data analytics with empirical software engineering techniques to rigorously re-evaluate key findings. The project advances a dynamic, reproducible, and sustainable measurement paradigm capable of adapting to evolving data ecosystems, offering a robust and timely framework for software metrics in the AI era.
This work addresses the inefficiency and limited scalability of manual evaluation of software engineering reproducibility packages by introducing, for the first time, a multi-agent architecture for assessing reproducibility quality. The authors translate open science guidelines into 31 machine-verifiable reproducibility criteria and integrate rule-based engines with automated scripts to perform both static and dynamic analyses of code, environments, and artifacts. The system generates evidence-based recommendations for improvement and supports human-in-the-loop optimization. Experimental evaluation on five reproducibility packages demonstrates 91.4% execution consistency and 75.4% detection accuracy, while user studies confirm its practical utility and strong potential for adoption.