🤖 AI Summary
This study addresses the irreproducibility of measurements in LLM inference benchmarking caused by execution and system state fluctuations. To mitigate this, we propose a sequential isolation methodology that suppresses inter-run noise through stepwise isolation protocols, while ensuring environmental reproducibility via an explicit cost breakeven model combined with Infrastructure as Code (IaC) practices. Evaluations of open-source models on NVIDIA A100 GPUs using the vLLM framework demonstrate that the proposed approach significantly reduces the mean coefficient of variation (CV) from 15.2% to 2.2%, with 78.5% of configurations achieving a CV below 3%. Furthermore, our experiments reveal critical latency mutation phenomena occurring between 200 and 500 concurrent users. These findings establish a rigorous, reproducible framework for LLM inference evaluation and highlight previously uncharacterized performance boundaries under high-concurrency workloads.
📝 Abstract
Reproducible benchmarking of Large Language Model (LLM) inference is challenging because repeated measurements can vary with execution and system state. We present the Sequential Isolation Methodology, a controlled benchmarking and regression-testing protocol designed to reduce between-run measurement variance while deliberately varying workload concurrency. We evaluate three representative open-source LLMs on an NVIDIA A100 80GB GPU using vLLM 0.9.1 across six context sizes and eight concurrency levels, with five repetitions per configuration. The final protocol reduces average coefficient of variation (CV) from 15.2% in the least controlled methodology stage to 2.2% under the final protocol; using CV computed across the five repetition-level median (P50) TTFT values per configuration, 113 of 144 configurations (78.5%) achieve CV below 3%. The measurements also show a marked latency transition between 200 and 500 concurrent users on the tested stack and descriptive differences in P99 latency across the three models. We additionally provide an explicit cost break-even model with sensitivity to API pricing. The protocol is intended to provide a stable reference for reproducible comparison and regression testing rather than to predict absolute behavior under uncontrolled production traffic. Infrastructure-as-Code and benchmark scripts support replication of the experimental environment.