SpecQuant: Speculative Decoding with Multi-Parent Quantization for Adaptive LLM Inference

📅 2026-09-18
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决本地运行大语言模型的计算和内存限制,SpecQuant结合推测解码与多父量化,自适应选择不同精度模型以提高推理效率。
📝 Abstract
Running large language models (LLMs) locally continues to be limited by restrictions of compute and memory on consumer hardware. The popular acceleration technologies, such as quantization, speculative decoding, and adaptive inferencing, offer substantial speed boosts but usually necessitate retraining, per architecture tuning, or draft models. SpecQuant is a trainingfree framework, that combines speculative decoding with multiparent quantization to perform adaptive, efficient inference of LLMs. SpecQuant derives multiple quantized variants (INT4, FP8, FP16) from a shared base model, and dynamically routes queries based on predicted complexity; lightweight variants are used for simple or factual tasks, and full-precision models are used for complex reasoning tasks or long-context inputs. The shared-weight design of SpecQuant ensures sufficient token acceptance for speculative decoding without compatibility issues using separate draft parent models. We evaluate SpecQuant on Qwen2.5 based models on the MMLU, AlpacaEval, and GSM8K datasets, or benchmarks, demonstrating 35-43% speedups without degrading accuracy greater than 2%, substantial within the LLM community. SpecQuant enables practical on-device LLM deployment across diverse hardware without special infrastructure or expertise.
Problem

Research questions and friction points this paper is trying to address.

large language models
consumer hardware
quantization
speculative decoding
adaptive inference
Innovation

Methods, ideas, or system contributions that make the work stand out.

speculative decoding
multi-parent quantization
adaptive inference
training-free framework
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Harish KB
Computer Science and Engineering, Vellore Institute of Technology, Vellore, India
J
Jagadeeswaran M
Computer Science and Engineering, Vellore Institute of Technology, Vellore, India
P
Pradheep P
Computer Science and Engineering, Vellore Institute of Technology, Vellore, India
Y
Yuvanesh S
Electronics and Communication Engineering, Vellore Institute of Technology, Vellore, India
S
Sivakumar T
School of Computer Science and Engineering, Vellore Institute of Technology, Vellore, India