Score
Designs and implements tests and benchmarks that measure how a software system’s performance (throughput, latency, resource utilization) changes as load and concurrency increase; builds and runs stress and scalability experiments, reproduces and documents results, analyzes bottlenecks, and guides changes to improve scalability.
This study addresses the high variability and low reliability of performance benchmarking results for stream-processing applications in cloud environments. Over three months, we conducted a large-scale longitudinal empirical study across multiple geographic regions and heterogeneous hardware—including diverse CPU architectures. Leveraging Kubernetes-based automated deployment, high-frequency repeated benchmarking, and time-series statistical analysis, we systematically characterized end-to-end cloud performance variability for the first time at the application level. We discovered that variability exhibits statistically significant diurnal and weekly periodicity (amplitude ≤2.5%) and a coefficient of variation <3.7%—substantially lower than commonly assumed in industry. Moreover, infrastructure sharing incurs at most a 2.5-percentage-point loss in measurement precision. These findings demonstrate strong robustness across regions and CPU architectures, providing empirical evidence and methodological foundations for enhancing reproducibility and trustworthiness in cloud-native benchmarking.
This work addresses the lack of systematic and rigorous performance benchmarking methodologies in programming language research, which has undermined the credibility of evaluation results. To remedy this, the paper introduces a closed-loop methodology—Measure-Explain-Test-Improve—that establishes, for the first time, a structured and reproducible workflow for performance assessment in the field. Integrating systematic experimental design, performance metric analysis, result interpretation, and iterative refinement, the approach emphasizes theoretical grounding and practical rigor at every stage. Its key contribution lies in enabling even researchers with limited empirical experience to conduct reliable and methodologically sound performance evaluations, thereby significantly enhancing the scientific validity and reproducibility of performance analysis in programming language research.
This study addresses the limited sensitivity of traditional cloud service performance regression detection, which is often hindered by I/O fluctuations and infrastructure changes. The authors propose a novel paradigm termed “Duet Instrumentation,” which uniquely integrates large language model (LLM)-driven code change analysis with synchronized dual-version benchmarking. By leveraging an LLM to precisely identify performance-relevant changes between consecutive versions, the method dynamically instruments only those critical code regions, achieving high-sensitivity regression detection with low overhead. Evaluated in real-world environments, the approach attains a precision of 58%, recall of 93%, and specificity of 71%, effectively detecting performance regressions as subtle as one-fifth the severity detectable by conventional methods.
Detecting performance regressions in configurable software is costly, and configuration sampling often misses localized performance degradation. Method: This paper proposes ConfFLARE, a technique that combines data-flow dependency analysis with change-impact propagation tracking to identify code changes interacting—via data flow—with performance-sensitive code. It further integrates configuration-feature identification to automatically select the subset of performance-sensitive configurations most likely affected by each change. Contribution/Results: ConfFLARE eliminates the need for exhaustive configuration-based performance testing. In evaluations on synthetic and real-world systems, it reduces the number of required test configurations by 79% and 70%, respectively, while achieving near-complete coverage of performance regression cases. It precisely pinpoints relevant features and significantly improves both the efficiency and completeness of performance regression detection.
To address the high computational cost and redundancy in benchmarking for software variant ranking, this paper proposes BISection Sampling (BISS), a correlation-aware test-suite reduction method that integrates critical-test preservation with a divide-and-conquer sampling strategy. BISS enables adaptive sampling across subsets of variants while preserving ranking stability. Evaluated on real-world datasets—including LLM leaderboards, SAT solver competitions, and configurable systems—BISS achieves an average 44% reduction in benchmarking overhead; in over 50% of cases, it reduces the number of required tests by up to 99%, without degrading Top-k ranking accuracy. The method thus provides a scalable, robust, and lightweight solution for large-scale variant assessment under resource constraints.
This work addresses the lack of existing tools capable of continuous performance validation and regression detection across entire datacenter clusters. The authors propose the first cluster-wide continuous benchmarking framework that supports unified scheduling, enabling simultaneous task distribution to all nodes and systematic collection of multidimensional performance metrics—spanning CPU, GPU, memory, interconnects, I/O, power consumption, frequency, and temperature—across both space and time. This framework facilitates performance regression detection under software and hardware changes as well as analysis of hardware variability. Experiments on the NHR@FAU cluster reveal intra-node performance variations below 1% among identically configured nodes, while inter-node differences reach up to 5%. The study further uncovers, for the first time, significant disparities in the performance–power relationship between air-cooled and liquid-cooled nodes.
Performance regression detection in large-scale software systems is hindered by the high overhead and low frequency of traditional benchmarking, limiting its integration into CI/CD pipelines. This paper introduces CloudBench, an efficient performance benchmarking platform designed for cloud-native CI/CD. Its core contributions are: (1) composable lightweight optimizations—including sampling, differential execution, and cache reuse—that drastically reduce benchmarking overhead; (2) an automated regression detection mechanism combining statistical hypothesis testing with SLA-aware thresholds; and (3) a highly available, declarative architecture enabling seamless integration with mainstream CI/CD toolchains. Experimental evaluation demonstrates that CloudBench achieves 99% detection accuracy while reducing average benchmark execution time by 7.3×, thereby enabling per-commit performance validation. To our knowledge, CloudBench provides the first production-ready, systematic solution for continuous performance engineering.
This study addresses the lack of systematic understanding regarding the energy impact of software refactoring under diverse workloads, particularly the absence of empirical analysis linking real-world refactoring practices to energy regressions. The authors construct a microbenchmark encompassing 68 refactoring types and a practical benchmark comprising 481 real refactoring commits from GitHub. Leveraging multi-workload scenarios and repeated paired energy measurements, they conduct the first large-scale investigation revealing that the energy effects of refactoring are highly workload-sensitive. Their findings demonstrate that refactoring type alone is insufficient to predict energy changes: 51.8% of refactorings in the microbenchmark and 7.5% in real projects induce statistically significant energy differences. Execution time explains energy variation only in controlled settings. Moreover, existing metrics and large language model–based approaches prove unreliable for detecting energy regressions.
This work addresses the limitation of existing test evolution benchmarks, which are confined to the method level and thus unable to automatically identify semantically stale or missing test cases at the project level. To bridge this gap, the authors introduce TEBench, the first benchmark specifically designed for project-level test evolution. It encompasses three scenarios—Test-Breaking, Test-Stale, and Test-Missing—and constructs task instances from Defects4J through a four-stage pipeline, providing developer-written ground-truth tests with fine-grained annotations. An evaluation across three agent frameworks and seven configurations involving six base models reveals that all approaches achieve only modest F1 scores (45.7%–49.4%) on test identification, with Test-Stale proving most challenging (F1 ≈ 36%). Moreover, while generated tests are often executable, they frequently deviate in form from the ground truth, exposing a fundamental limitation: current methods overly rely on execution-failure signals.
This work addresses the limitations of traditional high-performance computing (HPC), which relies on manual task scripting and scheduling and struggles to meet the automation demands of complex scientific workflows. The authors propose the first large language model–based autonomous agent framework that enables end-to-end automated execution of HPC workflows from descriptive instructions. The framework integrates Slurm/Flux job schedulers, low-latency AWS cloud infrastructure, and event monitoring mechanisms to support task definition, optimization, and scheduling. Experimental results demonstrate that the system efficiently deploys scalable experiments, accurately translates job specifications—with only occasional deviations in processor affinity—and successfully reproduces an expert-level variant calling pipeline, achieving consistent results in 18 out of 19 runs. These findings validate the framework’s feasibility and effectiveness in real-world HPC environments.