Score
Designs, implements, and validates benchmark workloads and workload generators — including event-based, synthesized, and modeled workloads — to exercise systems or models under controlled conditions. Measures and models workload characteristics, defines evaluation protocols and baselines, and computes performance metrics to compare systems or models across graded or constrained tasks.
Current evaluations of agent tool use often conflate workload specifications, action generation, and evidentiary criteria, lacking a unified and auditable framework. This work proposes an evaluation paradigm centered on “evidence admissibility gating,” which explicitly decouples workloads, drivers, and verification evidence through a shared evidence admissibility contract. The framework integrates diverse environments—including WebArena Verified, a subset of SWE-Gym, and MiniWoB++—and employs a standardized reporting pipeline comprising a universal workload adapter, declarative drivers, task manifests, event schemas, and replay/freeze strategies. It uniformly logs multidimensional metrics such as latency, invalid actions, and patching costs, enabling consistent differentiation of controller performance under identical workloads while ensuring relevance, reproducibility, and auditability in agent evaluations.
Existing benchmarks (e.g., TPC-H, TPC-DS) fail to capture distributional shifts and dynamic evolution characteristic of real production workloads, hindering the development and evaluation of learned database components. This paper introduces Redbench—the first high-fidelity benchmark built from 30 real-world cloud service production query workloads. Through systematic collection, statistical modeling, and pattern categorization, Redbench is the first to faithfully align with and reproduce the observed query distribution characteristics and temporal evolution patterns in Redset. It incorporates a workload alignment mechanism that ensures fidelity across critical dimensions: distributional shift, hotspot drift, and long-tail structure. Redbench supports reproducible query sampling and comprehensive workload feature analysis, significantly improving training effectiveness and generalization capability of learned components—such as indexes and query optimizers—in realistic deployment scenarios.
Current evaluations of large language model workflows lack proper calibration and fail to adequately reflect the severity of degradation. This work proposes WorkflowPerturb—the first benchmark framework based on controlled perturbations and severity grading—to systematically assess the sensitivity and calibration of multi-agent workflow metrics. By applying three types of perturbations—omission, compression, and description alteration—to 4,973 gold-standard workflows, we generate 44,757 perturbed variants. Through structured perturbation generation, multi-level perturbation control, and residual analysis of evaluation metrics, our approach reveals systematic differences among metric families, providing an interpretable and calibratable foundation for workflow evaluation.
To address low reusability of HPC benchmarks, poor cross-platform portability, and inefficient resource validation, this paper proposes the “benchmark carpentry” paradigm—a lightweight, reusable experimental execution framework. Methodologically, it integrates Cloudmesh’s experiment executor with HPE SmartSim, incorporating standardized workflow templates, AI/ML–simulation coupling mechanisms, and a unified experimental management interface. Its key contribution is the first application of craftsmanship principles to benchmarking process design, enabling automated, cross-domain and cross-architecture benchmark deployment and capability assessment. Evaluated on representative scientific computing workloads—including cloud masking analysis, seismic forecasting, and CFD surrogate modeling—the framework achieves ≥92% workflow reproducibility and reduces average deployment time by 68%, significantly improving resource configuration efficiency. It establishes a scalable, community-driven paradigm for HPC capability validation.
This paper systematically examines the structural role and evolutionary trajectory of simulation methods across the statistical lifecycle. Addressing the current fragmentation and conceptual ambiguity in simulation practice, the study introduces, for the first time, a comprehensive functional taxonomy—spanning model specification, diagnostic checking, validation, and inference—and proposes a “simulation-driven” paradigm for statistical practice, prioritizing computational scalability. Methodologically, it integrates Monte Carlo simulation, approximate Bayesian computation (ABC), simulation-based calibration, and posterior predictive checking, implemented via high-performance computing frameworks to enable large-scale empirical analysis. Key contributions are: (1) establishing simulation as foundational statistical infrastructure; (2) providing an actionable roadmap for algorithm design, statistical software development, and pedagogical reform; and (3) advancing a paradigm shift in statistical practice—from model-centric to simulation-augmented inference.
This work addresses the limitations of traditional high-performance computing (HPC), which relies on manual task scripting and scheduling and struggles to meet the automation demands of complex scientific workflows. The authors propose the first large language model–based autonomous agent framework that enables end-to-end automated execution of HPC workflows from descriptive instructions. The framework integrates Slurm/Flux job schedulers, low-latency AWS cloud infrastructure, and event monitoring mechanisms to support task definition, optimization, and scheduling. Experimental results demonstrate that the system efficiently deploys scalable experiments, accurately translates job specifications—with only occasional deviations in processor affinity—and successfully reproduces an expert-level variant calling pipeline, achieving consistent results in 18 out of 19 runs. These findings validate the framework’s feasibility and effectiveness in real-world HPC environments.
Existing benchmarks for knowledge work evaluation largely adhere to traditional NLP task paradigms, failing to capture systems’ capabilities in real-world knowledge-intensive settings. This work proposes a three-step framework—explicitly defining work activities, establishing realistic test environments, and focusing evaluation on deliverable outputs—and derives 18 core knowledge work activities from the O*NET database. Innovatively integrating role responsibilities, local tool usage, and downstream usability into benchmark design, the approach establishes a coherent “work activity–test setup–scoring artifact” alignment. Validation through three case studies (GDPval, OfficeQA Pro, and APEX-SWE) exposes critical misalignments in current benchmarks between tasks, environments, and actual work objectives, offering a new paradigm for evaluating knowledge work systems in practical, application-oriented contexts.
This study addresses the limitations of existing SysML verification approaches, which are often tool-dependent and restricted to performance properties, lacking support for automated validation of behavioral and interface requirements. To overcome these shortcomings, this work proposes a tool-agnostic, automated verification workflow driven by SysML test cases, integrating UML Testing Profile and behavioral diagram constructs to enable unified validation of multidimensional attributes—including behavior, timing, and state responses. The methodology was developed through a mixed-methods research strategy combining literature review and stakeholder interviews, and its efficacy was empirically validated across two independent SysML toolchains. The approach not only transcends the constraints of conventional parametric methods but also enables automatic traceability of verification results back to the original model elements.
Existing approaches struggle to effectively quantify the similarity and quality between synthetic and real data in evaluating tool-augmented agents. To address this gap, this work proposes SynAE, a novel framework that establishes the first multi-axis evaluation system tailored for multi-turn tool-use scenarios. SynAE introduces four fine-grained metric categories—assessing task instructions, tool invocations, final outputs, and downstream evaluation performance—to systematically measure synthetic data across dimensions of validity, fidelity, and diversity. Integrating natural language processing, trajectory modeling, and controllable generation techniques, the framework enables a reproducible evaluation pipeline and successfully identifies several representative failure modes in synthetic data generation. Empirical results demonstrate that such multidimensional assessment is essential for enhancing the reliability of agent evaluations.
Existing evaluation metrics struggle to assess the robustness of code agents in prolonged, multi-turn interactive programming scenarios. To address this gap, this work proposes the first black-box, language-agnostic benchmark centered on consecutive interaction rounds, driving agents to iteratively develop a REST API service through 100 programmatically generated change requests. The platform ensures reproducibility and realism by employing an isolated HTTP execution environment, a structured action space, programmatic testing, and a hybrid change sampler that emulates authentic “ambient programming” conditions. Experimental results reveal that all models fail within 5–6 rounds; however, incorporating a feedback-based retry mechanism improves success duration by up to 12×. Furthermore, high-performing agents exhibit significant sensitivity to the evaluation framework, with performance varying by as much as 6× between optimal and suboptimal configurations.