load testing

Designs, builds, and runs test harnesses and frameworks that generate, distribute, and measure workload to evaluate system and network performance, scalability, capacity, and reliability; this includes creating load-generation tools, distributed and automated load-testing platforms, and end-to-end or combined functional/performance scenarios. Uses workload assessment and benchmarking to validate performance, identify bottlenecks, guide tuning, and produce reproducible metrics and reports for performance validation and comparison.

loadtesting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
1.56
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$193K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Towards Easy and Realistic Network Infrastructure Testing for Large-scale Machine Learning

Apr 29, 2025
JY
Jinsun Yoo
🏛️ Georgia Institute of Technology | Harvard University | Hewlett Packard Labs

Evaluating network hardware behavior under large-scale ML training workloads is costly and low-fidelity due to reliance on expensive GPU-based testbeds. Method: This paper introduces Genie, a novel testing framework that generates realistic ML communication traffic solely on CPUs—eliminating the need for GPUs—and drives physical network devices on hardware testbeds; concurrently, it enhances ASTRA-sim to enable co-simulation of network microarchitectures and ML workloads. Contribution/Results: Genie establishes the first “CPU-driven + hardware-measured + simulation-enhanced” paradigm for network–ML co-evaluation, achieving, for the first time without GPUs, high-fidelity coupling between hardware-level network behavior and distributed training communication patterns. Across representative training scenarios, it achieves over 90% network performance prediction accuracy while reducing testing costs by 10×, significantly accelerating network architecture design and validation cycles.

Emulating GPU communication using CPU trafficModeling network-ML workload interaction via simulatorTesting network impact on ML performance without GPUs

This work addresses the challenges of SLO violations and resource inefficiency in machine learning model serving caused by inadequate capacity planning. To this end, the authors propose an adaptive, feedback-driven load testing framework that formalizes the ML serving load testing process for the first time. The framework incorporates real-traffic-based workload calibration and a warm-up mechanism, combined with adaptive search, performance signal feedback control, convergence detection, and GPU monitoring to efficiently estimate the maximum sustainable throughput under SLO constraints. Evaluation across 14 industrial cases demonstrates that the approach reduces capacity estimation error from approximately 30% to 2–6%, with the warm-up mechanism improving accuracy by 22.2%. This significantly mitigates deployment incidents and enhances GPU resource utilization efficiency.

capacity planningload testingML model serving

How to Evaluate Distributed Coordination Systems? -- A Survey and Analysis

Mar 14, 2024
BT
B. Turkkan
🏛️ IBM Research | University at Buffalo | University of New Hampshire | Microsoft | MongoDB

Existing distributed coordination services lack standardized testing methodologies and tools, resulting in incomplete evaluations and non-comparable results. Method: We conduct a systematic survey of evaluation practices across mainstream coordination services, identify critical gaps in benchmarking consistency, fault tolerance, and scalability, and distill six core evaluation requirements. Leveraging literature analysis and cross-tool comparison, we identify 12 key evaluation parameters and categorize five typical defects. Contribution/Results: We propose a standardized evaluation framework incorporating consistency models, fault injection, and distributed topology configurations. Furthermore, we introduce a reproducible, comparable, and scenario-driven benchmarking guideline—specifically designed for coordination services—that fills a critical gap in the domain’s dedicated evaluation ecosystem.

Coordination ServicesDistributed SystemsTesting Frameworks

Characterizing the Impact of Active Queue Management on Speed Test Measurements

Nov 24, 2025
SR
Siddhant Ray
🏛️ University of Chicago | Cal Poly | ENS Lyon

Existing speed measurement tools focus on peak throughput and poorly reflect users’ perceived responsiveness; emerging metrics such as “latency under load” show promise but their sensitivity to Active Queue Management (AQM) configurations remains unclear. Method: We empirically evaluate three mainstream AQM schemes—CoDel, FQ-CoDel, and SFQ—in a controlled network environment, systematically analyzing their impact on throughput and latency distributions, particularly latency under load. Results: AQM significantly alters speed test outcomes, with distinct latency-throughput trade-offs observed across algorithms under high load. Current measurement platforms, if uncalibrated for AQM, yield misleading latency estimates, undermining the reliability of policy and regulatory decisions. This study is the first to quantitatively characterize the structural impact of AQM on emerging speed metrics, providing critical empirical evidence to inform standardization of measurement tools and evidence-based network governance.

Calibrating speed tests for accurate policy guidanceComparing throughput variance across different AQM schemesUnderstanding AQM's impact on speed test latency metrics

Latest Papers

What's happening recently
View more

This work addresses the limitations of traditional high-performance computing (HPC), which relies on manual task scripting and scheduling and struggles to meet the automation demands of complex scientific workflows. The authors propose the first large language model–based autonomous agent framework that enables end-to-end automated execution of HPC workflows from descriptive instructions. The framework integrates Slurm/Flux job schedulers, low-latency AWS cloud infrastructure, and event monitoring mechanisms to support task definition, optimization, and scheduling. Experimental results demonstrate that the system efficiently deploys scalable experiments, accurately translates job specifications—with only occasional deviations in processor affinity—and successfully reproduces an expert-level variant calling pipeline, achieving consistent results in 18 out of 19 runs. These findings validate the framework’s feasibility and effectiveness in real-world HPC environments.

Autonomous AgentsHigh Performance ComputingJob Specification Translation

This work addresses the lack of existing tools capable of continuous performance validation and regression detection across entire datacenter clusters. The authors propose the first cluster-wide continuous benchmarking framework that supports unified scheduling, enabling simultaneous task distribution to all nodes and systematic collection of multidimensional performance metrics—spanning CPU, GPU, memory, interconnects, I/O, power consumption, frequency, and temperature—across both space and time. This framework facilitates performance regression detection under software and hardware changes as well as analysis of hardware variability. Experiments on the NHR@FAU cluster reveal intra-node performance variations below 1% among identically configured nodes, while inter-node differences reach up to 5%. The study further uncovers, for the first time, significant disparities in the performance–power relationship between air-cooled and liquid-cooled nodes.

cluster-wide benchmarkingcontinuous testingdata center validation

Existing benchmarks struggle to evaluate the impact of harness components on the performance of large language model (LLM) agent systems, often overlooking execution details or fixing harness configurations. This work proposes Harness-Bench, the first framework to systematically decouple and quantify how harness design influences agent workflows. By enforcing a unified task environment, computational budget, and evaluation protocol—and leveraging sandboxed offline tasks, realistic usage patterns, human auditing, and full execution trace logging—it enables controlled experiments across diverse model–harness combinations. Analysis of 5,194 execution traces across 106 tasks reveals significant differences in completion rates, process quality, efficiency, and failure modes, exposing alignment failures stemming from misalignment between reasoning and execution. The study argues that agent performance should be reported based on joint model–harness configurations rather than base models alone.

agent workflowsbenchmarkingexecution-layer variation

Current evaluations of large models predominantly rely on end-to-end metrics, which obscure the underlying causes of performance variations due to hardware and software configurations. This work proposes the first reproducible, execution-trace-based benchmarking framework that constructs a community-extensible, trace-level evidence ecosystem through fine-grained execution traces, YAML-based workload specifications, and containerized launch scripts. The framework enables in-depth analysis of computational, memory, and communication efficiency. Using this approach, the study systematically quantifies—for the first time—the impact of parallelization strategies, interconnect bandwidth, and framework-level optimizations on training performance. Key findings include: high compute-communication overlap does not necessarily reduce step time; doubling TPU interconnect bandwidth yields significantly greater benefits than on GPUs for small-to-medium workloads; and performance gaps of up to 3× exist between optimal configurations across different frameworks.

benchmarkingconfiguration spaceLLM infrastructure

This work addresses the lack of systematic, reproducible, and maintainable testing methodologies in existing dynamic resource management libraries. We propose an automated validation framework tailored for high-performance computing (HPC) environments, which introduces a novel multi-level testing taxonomy encompassing both functional and non-functional requirements. Built upon an MPI-based scalable library testing methodology, the framework supports core primitives of dynamic resource management systems—such as initialization, readiness checks, and reconfiguration—and integrates containerized virtual clusters with continuous integration (CI) ecosystems. Experimental evaluation demonstrates that our approach significantly improves early fault detection rates, reduces maintenance overhead caused by evolving dependencies, and is readily generalizable to other systems exhibiting similar variability mechanisms.

dynamic resource managementHPClibrary correctness

Hot Scholars

EV

Emiliano Votta

Politecnico di Milano
BiomechanicsHeart Valve SurgeryCardiac Care
MX

Minxian Xu

Associate Professor, Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences
Cloud ComputingMicroservicesLLM Inference
KY

Kejiang Ye

Professor, Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences
Cloud ComputingAI SystemsIndustrial Internet
JR

Ji-Rong Wen

Gaoling School of Artificial Intelligence, Renmin University of China
Large Language ModelWeb SearchInformation RetrievalMachine Learning
XB

Xiaohe Bo

Gaoling School of Artificial Intelligence, Renmin University of China
large language models