Reproducible LLM Inference Benchmarking: A Sequential Isolation Protocol for Regression Testing

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the irreproducibility of measurements in LLM inference benchmarking caused by execution and system state fluctuations. To mitigate this, we propose a sequential isolation methodology that suppresses inter-run noise through stepwise isolation protocols, while ensuring environmental reproducibility via an explicit cost breakeven model combined with Infrastructure as Code (IaC) practices. Evaluations of open-source models on NVIDIA A100 GPUs using the vLLM framework demonstrate that the proposed approach significantly reduces the mean coefficient of variation (CV) from 15.2% to 2.2%, with 78.5% of configurations achieving a CV below 3%. Furthermore, our experiments reveal critical latency mutation phenomena occurring between 200 and 500 concurrent users. These findings establish a rigorous, reproducible framework for LLM inference evaluation and highlight previously uncharacterized performance boundaries under high-concurrency workloads.
📝 Abstract
Reproducible benchmarking of Large Language Model (LLM) inference is challenging because repeated measurements can vary with execution and system state. We present the Sequential Isolation Methodology, a controlled benchmarking and regression-testing protocol designed to reduce between-run measurement variance while deliberately varying workload concurrency. We evaluate three representative open-source LLMs on an NVIDIA A100 80GB GPU using vLLM 0.9.1 across six context sizes and eight concurrency levels, with five repetitions per configuration. The final protocol reduces average coefficient of variation (CV) from 15.2% in the least controlled methodology stage to 2.2% under the final protocol; using CV computed across the five repetition-level median (P50) TTFT values per configuration, 113 of 144 configurations (78.5%) achieve CV below 3%. The measurements also show a marked latency transition between 200 and 500 concurrent users on the tested stack and descriptive differences in P99 latency across the three models. We additionally provide an explicit cost break-even model with sensitivity to API pricing. The protocol is intended to provide a stable reference for reproducible comparison and regression testing rather than to predict absolute behavior under uncontrolled production traffic. Infrastructure-as-Code and benchmark scripts support replication of the experimental environment.
Problem

Research questions and friction points this paper is trying to address.

LLM inference benchmarking
reproducibility
regression testing
measurement variance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sequential Isolation Methodology
Reproducible Benchmarking
Regression Testing
Coefficient of Variation
Infrastructure-as-Code
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Arnold Olympio
Lucerne University of Applied Sciences and Arts, Switzerland
J
Juan Manuel Servera Bondroit
Microsoft, Switzerland
W
Wael Abdelmalek
Uthereal AG, Switzerland
Guang Lu
Guang Lu
Lucerne University of Applied Sciences and Arts, Switzerland
João Carvalho
João Carvalho
Technische Universität Darmstadt
Machine LearningRoboticsReinforcement Learning