perform scalability testing

Designs and implements tests and benchmarks that measure how a software system’s performance (throughput, latency, resource utilization) changes as load and concurrency increase; builds and runs stress and scalability experiments, reproduces and documents results, analyzes bottlenecks, and guides changes to improve scalability.

performscalabilitytesting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.3
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$205K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

When Should I Run My Application Benchmark?: Studying Cloud Performance Variability for the Case of Stream Processing Applications

Apr 16, 2025
SH
Soren Henning
🏛️ Dynatrace Research | Johannes Kepler University Linz

This study addresses the high variability and low reliability of performance benchmarking results for stream-processing applications in cloud environments. Over three months, we conducted a large-scale longitudinal empirical study across multiple geographic regions and heterogeneous hardware—including diverse CPU architectures. Leveraging Kubernetes-based automated deployment, high-frequency repeated benchmarking, and time-series statistical analysis, we systematically characterized end-to-end cloud performance variability for the first time at the application level. We discovered that variability exhibits statistically significant diurnal and weekly periodicity (amplitude ≤2.5%) and a coefficient of variation <3.7%—substantially lower than commonly assumed in industry. Moreover, infrastructure sharing incurs at most a 2.5-percentage-point loss in measurement precision. These findings demonstrate strong robustness across regions and CPU architectures, providing empirical evidence and methodological foundations for enhancing reproducibility and trustworthiness in cloud-native benchmarking.

Assess temporal effects on stream processing applicationsEvaluate benchmark result accuracy across cloud regionsQuantify cloud performance variability impact on benchmarks

This work addresses the lack of systematic and rigorous performance benchmarking methodologies in programming language research, which has undermined the credibility of evaluation results. To remedy this, the paper introduces a closed-loop methodology—Measure-Explain-Test-Improve—that establishes, for the first time, a structured and reproducible workflow for performance assessment in the field. Integrating systematic experimental design, performance metric analysis, result interpretation, and iterative refinement, the approach emphasizes theoretical grounding and practical rigor at every stage. Its key contribution lies in enabling even researchers with limited empirical experience to conduct reliable and methodologically sound performance evaluations, thereby significantly enhancing the scientific validity and reproducibility of performance analysis in programming language research.

benchmarkingperformance evaluationprogramming language research

This study addresses the limited sensitivity of traditional cloud service performance regression detection, which is often hindered by I/O fluctuations and infrastructure changes. The authors propose a novel paradigm termed “Duet Instrumentation,” which uniquely integrates large language model (LLM)-driven code change analysis with synchronized dual-version benchmarking. By leveraging an LLM to precisely identify performance-relevant changes between consecutive versions, the method dynamically instruments only those critical code regions, achieving high-sensitivity regression detection with low overhead. Evaluated in real-world environments, the approach attains a precision of 58%, recall of 93%, and specificity of 71%, effectively detecting performance regressions as subtle as one-fifth the severity detectable by conventional methods.

application benchmarkscloud service benchmarkingmicrobenchmarks

Detecting Performance-Relevant Changes in Configurable Software Systems

Nov 21, 2025
SB
Sebastian Böhm
🏛️ Saarland University | Leipzig University

Detecting performance regressions in configurable software is costly, and configuration sampling often misses localized performance degradation. Method: This paper proposes ConfFLARE, a technique that combines data-flow dependency analysis with change-impact propagation tracking to identify code changes interacting—via data flow—with performance-sensitive code. It further integrates configuration-feature identification to automatically select the subset of performance-sensitive configurations most likely affected by each change. Contribution/Results: ConfFLARE eliminates the need for exhaustive configuration-based performance testing. In evaluations on synthetic and real-world systems, it reduces the number of required test configurations by 79% and 70%, respectively, while achieving near-complete coverage of performance regression cases. It precisely pinpoints relevant features and significantly improves both the efficiency and completeness of performance regression detection.

Detecting performance-relevant changes in configurable software systems efficientlyIdentifying performance regressions through data-flow interactions with critical codeReducing measurement costs by selecting relevant configurations for testing

Efficiently Ranking Software Variants with Minimal Benchmarks

Sep 08, 2025
TM
Théo Matricon
🏛️ Univ Rennes Inria | CNRS | IRISA | IUF | Simula Research Laboratory

To address the high computational cost and redundancy in benchmarking for software variant ranking, this paper proposes BISection Sampling (BISS), a correlation-aware test-suite reduction method that integrates critical-test preservation with a divide-and-conquer sampling strategy. BISS enables adaptive sampling across subsets of variants while preserving ranking stability. Evaluated on real-world datasets—including LLM leaderboards, SAT solver competitions, and configurable systems—BISS achieves an average 44% reduction in benchmarking overhead; in over 50% of cases, it reduces the number of required tests by up to 99%, without degrading Top-k ranking accuracy. The method thus provides a scalable, robust, and lightweight solution for large-scale variant assessment under resource constraints.

Maintaining stable rankings with fewer testsOptimizing test suites to minimize computational resourcesReducing benchmark costs for software variants

Latest Papers

What's happening recently
View more

This work addresses the lack of existing tools capable of continuous performance validation and regression detection across entire datacenter clusters. The authors propose the first cluster-wide continuous benchmarking framework that supports unified scheduling, enabling simultaneous task distribution to all nodes and systematic collection of multidimensional performance metrics—spanning CPU, GPU, memory, interconnects, I/O, power consumption, frequency, and temperature—across both space and time. This framework facilitates performance regression detection under software and hardware changes as well as analysis of hardware variability. Experiments on the NHR@FAU cluster reveal intra-node performance variations below 1% among identically configured nodes, while inter-node differences reach up to 5%. The study further uncovers, for the first time, significant disparities in the performance–power relationship between air-cooled and liquid-cooled nodes.

cluster-wide benchmarkingcontinuous testingdata center validation

Towards an Optimized Benchmarking Platform for CI/CD Pipelines

Oct 21, 2025
NJ
Nils Japke
🏛️ Technische Universität Berlin | DATEV eG

Performance regression detection in large-scale software systems is hindered by the high overhead and low frequency of traditional benchmarking, limiting its integration into CI/CD pipelines. This paper introduces CloudBench, an efficient performance benchmarking platform designed for cloud-native CI/CD. Its core contributions are: (1) composable lightweight optimizations—including sampling, differential execution, and cache reuse—that drastically reduce benchmarking overhead; (2) an automated regression detection mechanism combining statistical hypothesis testing with SLA-aware thresholds; and (3) a highly available, declarative architecture enabling seamless integration with mainstream CI/CD toolchains. Experimental evaluation demonstrates that CloudBench achieves 99% detection accuracy while reducing average benchmark execution time by 7.3×, thereby enabling per-commit performance validation. To our knowledge, CloudBench provides the first production-ready, systematic solution for continuous performance engineering.

Detecting performance regressions in CI/CD pipelines earlyIntegrating benchmark optimizations into practical CI/CD systemsOptimizing resource-intensive benchmarks for efficient execution

This study addresses the lack of systematic understanding regarding the energy impact of software refactoring under diverse workloads, particularly the absence of empirical analysis linking real-world refactoring practices to energy regressions. The authors construct a microbenchmark encompassing 68 refactoring types and a practical benchmark comprising 481 real refactoring commits from GitHub. Leveraging multi-workload scenarios and repeated paired energy measurements, they conduct the first large-scale investigation revealing that the energy effects of refactoring are highly workload-sensitive. Their findings demonstrate that refactoring type alone is insufficient to predict energy changes: 51.8% of refactorings in the microbenchmark and 7.5% in real projects induce statistically significant energy differences. Execution time explains energy variation only in controlled settings. Moreover, existing metrics and large language model–based approaches prove unreliable for detecting energy regressions.

energy impactenergy regressionreal-world commits

This work addresses the limitation of existing test evolution benchmarks, which are confined to the method level and thus unable to automatically identify semantically stale or missing test cases at the project level. To bridge this gap, the authors introduce TEBench, the first benchmark specifically designed for project-level test evolution. It encompasses three scenarios—Test-Breaking, Test-Stale, and Test-Missing—and constructs task instances from Defects4J through a four-stage pipeline, providing developer-written ground-truth tests with fine-grained annotations. An evaluation across three agent frameworks and seven configurations involving six base models reveals that all approaches achieve only modest F1 scores (45.7%–49.4%) on test identification, with Test-Stale proving most challenging (F1 ≈ 36%). Moreover, while generated tests are often executable, they frequently deviate in form from the ground truth, exposing a fundamental limitation: current methods overly rely on execution-failure signals.

code evolutionproject-level testingsoftware testing

This work addresses the limitations of traditional high-performance computing (HPC), which relies on manual task scripting and scheduling and struggles to meet the automation demands of complex scientific workflows. The authors propose the first large language model–based autonomous agent framework that enables end-to-end automated execution of HPC workflows from descriptive instructions. The framework integrates Slurm/Flux job schedulers, low-latency AWS cloud infrastructure, and event monitoring mechanisms to support task definition, optimization, and scheduling. Experimental results demonstrate that the system efficiently deploys scalable experiments, accurately translates job specifications—with only occasional deviations in processor affinity—and successfully reproduces an expert-level variant calling pipeline, achieving consistent results in 18 out of 19 runs. These findings validate the framework’s feasibility and effectiveness in real-world HPC environments.

Autonomous AgentsHigh Performance ComputingJob Specification Translation

Hot Scholars

HX

Hui Xiong

Senior Scientist, Candela Corporation
Ultrafast dynamicsatomic molecular physicsfree electron laser
RH

Rong-Hua Li

Beijing Institute of Technology
Algorithms for (big) graphmatrixand sequence data
CX

Caiming Xiong

Salesforce Research
Machine LearningNLPComputer VisionMultimedia
SH

Shelby Heinecke

Salesforce Research
Artificial IntelligenceAI AgentsLLM AgentsMulti-Agent Systems
SS

Silvio Savarese

Associate Professor of Computer Science at Stanford University
Computer vision