π€ AI Summary
This study addresses the absence of formal specifications for deterministic outputs in existing LLM inference systems, which leads to inconsistent results across execution variants. This work establishes the first system-level specification for deterministic inference and proposes Vosti, an engine that enforces strict consistency through independent kernel selection and logically prefix-bound KV caches. We leverage Verus to construct inductive proofs of scheduling and caching logic at the engineβGPU kernel boundary, complemented by a Triton analyzer that guarantees bitwise consistency across batches and memory layouts. Experimental evaluations demonstrate that Vosti achieves bitwise-identical logits across all tested variants while delivering performance comparable to vLLM on decode-intensive workloads. Ultimately, this work provides strong formal guarantees suitable for production deployment.
π Abstract
LLM inference systems may vary batch composition, prompt chunking, prefill/decode execution, and KV-cache reuse, eviction, or recomputation. These optimizations should not affect system outputs. Production systems, including vLLM's batch-invariant mode and SGLang's deterministic mode, target this goal but lack a formal system-level specification.
We formalize deterministic LLM inference: under a fixed model and deployment configuration, requests with the same prompt and initial sampler state produce bitwise-identical logits at corresponding output positions across executions. Our tests find that these production modes produce different logits under some execution variations. To address this limitation, we present Vosti, an inference engine designed and verified against this specification. Vosti chooses kernels independently of runtime engine state and ties cached KV values to their logical token prefixes. Its proof decomposes at the engine/GPU kernel boundary: a Verus inductive proof establishes that scheduling and the paged, prefix-sharing KV-cache preserve output logits, while a Triton analyzer proves bitwise-equal selected kernel outputs across batches, query lengths, and paged KV-cache layouts. Vosti produces bitwise-identical logits across every tested execution variation and achieves performance comparable to vLLM's batch-invariant mode on decode-heavy workloads, while providing a stronger, formally verified determinism guarantee.