Spexis: Speculative Lookahead Scheduling for LLM Inference

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the efficiency bottlenecks and memory pressure associated with pipeline and tensor parallelism in multi-GPU large language model inference by proposing Spexis, a framework built upon vLLM. Spexis innovatively introduces a speculative parallel dimension that decouples speculative execution from standard inference, enabling their concurrent operation. Furthermore, it designs a lookahead scheduling mechanism that predicts generation quality to minimize futile computation and unnecessary KV cache eviction, achieving efficient coordination with zero additional memory overhead. Experimental results demonstrate that Spexis significantly improves system memory efficiency, delivering up to a 34% increase in serving throughput over state-of-the-art baselines across diverse multi-GPU configurations.
📝 Abstract
Spexis is a multi-GPU LLM inference framework that improves the efficiency of pipeline and tensor parallelism through speculative parallelism. Rather than using speculative decoding only to accelerate token generation, Spexis runs speculation in parallel with normal execution, introducing a new parallelism axis without increasing KV-cache memory usage. This improves memory efficiency and helps mitigate the bottlenecks of multi-GPU inference. Spexis further uses lookahead scheduling to predict speculation quality and future memory pressure, allowing it to reduce wasted speculation, KV-cache eviction, and recomputation. Built on top of vLLM, Spexis largely improves serving performance across a range of GPU configurations, achieving speedups of up to 34% over a baseline that uses the optimal combination of pipeline and tensor parallelism. Spexis's source code is publicly available at https://github.com/mlsys-seo/spexis.
Problem

Research questions and friction points this paper is trying to address.

LLM inference
multi-GPU
pipeline parallelism
tensor parallelism
KV-cache memory
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speculative Parallelism
Lookahead Scheduling
Multi-GPU Inference
Pipeline Parallelism
Tensor Parallelism
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.