arXiv Science⌕ Search

arXiv · 2610.09778

Reproducible LLM Inference Benchmarking: A Sequential Isolation Protocol for Regression Testing

Abstract

Reproducible benchmarking of Large Language Model (LLM) inference is challenging because repeated measurements can vary with execution and system state. We present the Sequential Isolation Methodology, a controlled benchmarking and regression-testing protocol designed to reduce between-run measurement variance while deliberately varying workload concurrency. We evaluate three representative open-source LLMs on an NVIDIA A100 80GB GPU using vLLM 0.9.1 across six context sizes and eight concurrency levels, with five repetitions per configuration. The final protocol reduces average coefficient of variation (CV) from 15.2% in the least controlled methodology stage to 2.2% under the final protocol; using CV computed across the five repetition-level median (P50) TTFT values per configuration, 113 of 144 configurations (78.5%) achieve CV below 3%. The measurements also show a marked latency transition between 200 and 500 concurrent users on the tested stack and descriptive differences in P99 latency across the three models. We additionally provide an explicit cost break-even model with sensitivity to API pricing. The protocol is intended to provide a stable reference for reproducible comparison and regression testing rather than to predict absolute behavior under uncontrolled production traffic. Infrastructure-as-Code and benchmark scripts support replication of the experimental environment.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Arnold Olympio, Juan Manuel Servera Bondroit, Wael Abdelmalek, Guang Lu, João Carvalho. 2026-10-07. Reproducible LLM Inference Benchmarking: A Sequential Isolation Protocol for Regression Testing. https://arxiv.org/abs/2610.09778

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

CXDVirt: Low Latency Kernel Module Based CXL-SSD Emulation

CXL-SSDs promise memory-semantic access to NAND-scale capacity through a small in-device DRAM cache, but scarce hardware makes emulation central to CXL-SSD research. Cylon, a state-of-the-art emulator, runs workloads in a VM, intercepting DRAM misses to inject NAND latency. We show this VM machinery adds over 3$μs$ per miss, exceeding high-performance NAND read latency, while repeated VM exits inflate tail miss latencies several-fold. We present CXDVirt, a kernel-module-based CXL-SSD emulator that removes the VM entirely: the device is exposed through host devdax, hits run natively via the MMU, and a custom page-fault path models NAND latency on misses. CXDVirt preserves CXL-SSDs' bimodal latency profile while reducing per-miss latency 1.4-1.6$\times$ (4.1$\times$ at P99.9); with eight threads its miss latency only doubles, against 8-9$\times$ for Cylon behind QEMU's global lock. CXDVirt stays within 5% of remote DRAM without NAND traffic, where Cylon is 3.1-3.6$\times$ slower, and enables configurable eviction and prefetch policy studies.

cs.PF↗

LLBPE: Linked-List Based GPU-Parallel BPE Tokenizer

Every LLM inference begins with tokenization, which converts raw input bytes into the discrete token sequence the model consumes. For text, this step is often implemented using Byte Pair Encoding (BPE), an algorithm originally introduced for data compression. BPE has traditionally run on the CPU with extensive optimization, but recent work has moved it to the GPU for higher throughput. We show that these GPU implementations are bottlenecked not by computation but by data movement. We develop LLBPE that represents the token sequence as an array-based linked list so that each merge reduces to a constant- time pointer update. Furthermore, LLBPE fuses rank lookup, minimum selection, and merging into a single kernel to eliminate redundant hash map queries. LLBPE achieves up to 5.2x higher throughput than the best existing GPU implementation and 24.6x over optimized CPU implementations, at the cost of minor discrepancies in tokenized output.

cs.PF↗

Energy-performance tradeoffs in server farms with batch services and setup times

Data centers consume a large amount of energy, much of which is wasted due to idle servers. Turning off idle servers might be an effective power-saving solution; however, there is a trade-off between energy savings and system performance. Hence, we propose a setup queueing model with a batching policy that allows servers to process a set of jobs simultaneously to minimize power consumption while maintaining acceptable performance. We consider an M/M/c/SET--BATCH queue, a multi-server batch service queue with a fixed batch size and setup times, and some variants, including systems in which idle servers delay before turning off or systems in which the batch size is dynamic. We analyze the steady-state probabilities and system performance of the M/M/c/SET--BATCH system and its variants. Our analysis of the M/M/c/SET--BATCH system with lower computational complexity is made possible by utilizing the special structure of the model. In addition, we use simulations to compare the M/M/c/SET--BATCH model with some other variants with different setup time distributions. The results suggest that the model performs better when the setup time has a larger coefficient of variation. Our results indicate that the batching policy enhances the system performance, especially when we allow servers to be idle before turning them off.

cs.PF↗