build benchmarking pipelines

Design and implement end-to-end systems that construct, manage, and run reproducible benchmarking pipelines: standardize preprocessing, generate and freeze train/validation/held-out splits, assemble dynamic or hybrid benchmark instances, and emit artifacts (snapshots, metric records, and replayable views) for diagnostics and transfer evaluation. This includes automated and large-scale processes for benchmark generation and filtering (e.g., perplexity‑consistency or other selection), creation of mosaic or bugfix-focused datasets, and protocol‑compliant workflow engineering for shareable, replicable evaluation.

buildbenchmarkingpipelines

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.55
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$224K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the reproducibility challenges posed by the rapid evolution of large models and high-performance computing systems, where existing benchmarks lack sustainable and automated evaluation mechanisms. To bridge this gap, the authors propose a user-agnostic continuous benchmarking framework that integrates principles from software engineering—particularly continuous integration—to establish an automated pipeline. This pipeline seamlessly combines systematic workflows with community-driven collaboration, delivering a reproducible and scalable benchmarking infrastructure for artificial intelligence and neuroscience research. The framework significantly enhances the sustainability, transparency, and collaborative efficiency of scientific evaluation in these fields.

automated benchmarkingcontinuous benchmarkinghigh-performance computing

When Should I Run My Application Benchmark?: Studying Cloud Performance Variability for the Case of Stream Processing Applications

Apr 16, 2025
SH
Soren Henning
🏛️ Dynatrace Research | Johannes Kepler University Linz

This study addresses the high variability and low reliability of performance benchmarking results for stream-processing applications in cloud environments. Over three months, we conducted a large-scale longitudinal empirical study across multiple geographic regions and heterogeneous hardware—including diverse CPU architectures. Leveraging Kubernetes-based automated deployment, high-frequency repeated benchmarking, and time-series statistical analysis, we systematically characterized end-to-end cloud performance variability for the first time at the application level. We discovered that variability exhibits statistically significant diurnal and weekly periodicity (amplitude ≤2.5%) and a coefficient of variation <3.7%—substantially lower than commonly assumed in industry. Moreover, infrastructure sharing incurs at most a 2.5-percentage-point loss in measurement precision. These findings demonstrate strong robustness across regions and CPU architectures, providing empirical evidence and methodological foundations for enhancing reproducibility and trustworthiness in cloud-native benchmarking.

Assess temporal effects on stream processing applicationsEvaluate benchmark result accuracy across cloud regionsQuantify cloud performance variability impact on benchmarks

Towards an Optimized Benchmarking Platform for CI/CD Pipelines

Oct 21, 2025
NJ
Nils Japke
🏛️ Technische Universität Berlin | DATEV eG

Performance regression detection in large-scale software systems is hindered by the high overhead and low frequency of traditional benchmarking, limiting its integration into CI/CD pipelines. This paper introduces CloudBench, an efficient performance benchmarking platform designed for cloud-native CI/CD. Its core contributions are: (1) composable lightweight optimizations—including sampling, differential execution, and cache reuse—that drastically reduce benchmarking overhead; (2) an automated regression detection mechanism combining statistical hypothesis testing with SLA-aware thresholds; and (3) a highly available, declarative architecture enabling seamless integration with mainstream CI/CD toolchains. Experimental evaluation demonstrates that CloudBench achieves 99% detection accuracy while reducing average benchmark execution time by 7.3×, thereby enabling per-commit performance validation. To our knowledge, CloudBench provides the first production-ready, systematic solution for continuous performance engineering.

Detecting performance regressions in CI/CD pipelines earlyIntegrating benchmark optimizations into practical CI/CD systemsOptimizing resource-intensive benchmarks for efficient execution

This work addresses the inefficiency in notebook-based distributed workflows, where minor modifications often trigger full re-execution, severely hindering iterative development and reproducibility. To overcome this limitation, the authors propose NBRewind, a system that, for the first time, enables fine-grained incremental execution and cross-platform portability while preserving reproducibility. NBRewind integrates a dual-kernel architecture—comprising auditing and replay components—with cell-level incremental checkpoints and inter-cell dataflow analysis. It further leverages standardized notebook packaging to facilitate efficient partial re-execution. Evaluation in real-world high-performance computing (HPC) scenarios demonstrates that NBRewind incurs minimal overhead for incremental checkpointing and substantially improves both execution efficiency and cross-site reproducibility.

checkpointingdistributed workflowsiterative development

Existing code generation benchmarks are largely confined to function-level tasks or localized modifications, lacking the capacity to evaluate end-to-end generation of complete microservice repositories from scratch. This work proposes RepoGenesis, the first benchmark for multilingual, end-to-end microservice repository generation, encompassing 106 real-world Python and Java projects spanning 11 frameworks and 18 domains. Data quality is ensured through a rigorous "review-rebuttal" curation process, and novel evaluation metrics—including API coverage and deployment success rate—are introduced. Experimental results show that the best-performing system achieves Pass@1 scores of 23.67% and 21.45% on Python and Java, respectively, with a deployment success rate of up to 100%, though challenges remain in architectural coherence. Notably, GenesisAgent-8B, fine-tuned on this benchmark, matches the performance of GPT-5 mini.

end-to-end developmentLLM benchmarkingmicroservice generation

Latest Papers

What's happening recently
View more

Traditional benchmarks provide only aggregate scores, offering insufficient evidence to support reliable deployment decisions and thereby creating a disconnect between evaluation and action. To address this gap, this work proposes a “deployment-completeness” benchmarking framework, introducing novel metrics—evidence fibers, completeness curves, and certifiable proportions—alongside a systematic audit methodology comprising evidence fiber analysis, response ranking intervals, conformal coverage evaluation, and a certify-then-acquire decision pipeline. Empirical evaluation on benchmarks such as Tox21, Matbench, and JARVIS reveals that conventional approaches suffer a drastic drop in channel coverage to 10.07% under real-world deployment conditions. In contrast, the proposed method reduces error-driven deployment decisions to 0.027% on Tox21 and 0.128% on JARVIS, substantially enhancing deployment reliability.

benchmark evidencecertifiable fractiondeployment action

Existing program repair benchmarks inadequately reflect real-world repository-level continuous integration (CI) scenarios, as they overlook critical challenges such as non-code artifacts, environmental dependencies, and workflow constraints. This work introduces the first repository-level repair benchmark grounded in actual GitHub Actions executions, validating patches through faithful replay of original CI workflows. The benchmark includes 567 CI failures meticulously annotated into 12 fine-grained error categories. Innovatively adopting end-to-end CI workflow re-execution as the patch validation criterion, it enables error-type-aware evaluation. By integrating log analysis, fault localization, and large language model–generated candidate patches, the approach achieves strong performance on tool-enforced errors like formatting and static checks, attaining an overall best repair success rate of 18.9%, while environment- and configuration-related issues remain notably challenging.

Automated Patch ValidationCI FailuresContinuous Integration

Current evaluations of agent tool use often conflate workload specifications, action generation, and evidentiary criteria, lacking a unified and auditable framework. This work proposes an evaluation paradigm centered on “evidence admissibility gating,” which explicitly decouples workloads, drivers, and verification evidence through a shared evidence admissibility contract. The framework integrates diverse environments—including WebArena Verified, a subset of SWE-Gym, and MiniWoB++—and employs a standardized reporting pipeline comprising a universal workload adapter, declarative drivers, task manifests, event schemas, and replay/freeze strategies. It uniformly logs multidimensional metrics such as latency, invalid actions, and patching costs, enabling consistent differentiation of controller performance under identical workloads while ensuring relevance, reproducibility, and auditability in agent evaluations.

benchmarkingevaluation methodologyevidence admission

Community-driven scientific workflow ecosystems often struggle to sustain themselves due to ambiguous maintenance and user support mechanisms, particularly in cross-platform collaboration and heterogeneous execution environments. This study presents the first cross-platform empirical analysis of the nf-core ecosystem, systematically examining 15,760 GitHub issues, 35,411 pull requests, and 895 forum discussions. By integrating metadata and textual features into predictive models, the research uncovers significant disparities in maintenance and support activities across platforms and highlights weak explicit linkages among them. The findings reveal that issues, pull requests, and forum posts predominantly serve distinct roles—coordinating maintenance, facilitating code integration, and providing user support, respectively. Moreover, issue actionability, diagnostic evidence, and depth of interaction emerge as critical determinants of resolution efficiency.

community-drivenheterogeneous execution environmentsmaintenance

Hot Scholars

GZ

Guangtao Zhai

Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI Evaluation
MB

Mohit Bansal

Parker Distinguished Professor, Computer Science, UNC Chapel Hill
Natural Language ProcessingComputer VisionMachine LearningMultimodal AI
HW

Haoning Wu

Shanghai Jiao Tong University
Computer VisionMulti-modal LearningGenerative Models