Disaggregated Quantization: Specializing LLM Prefill and Decode

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为了解决预填充和解码阶段的量化问题,提出了一种名为'分散量化'(DQ)的方法,通过专门化计算格式、权重及存储位置来优化这两个阶段,从而在不增加推理成本的情况下提高准确性。
📝 Abstract
Prefill and decode reward different approaches to quantization: low-precision arithmetic accelerates prompt processing, while compact weights reduce memory traffic during generation. We propose "disaggregated quantization" (DQ), which specializes computation formats, weights and storage placement to both of these phases. On Qwen 3 and Gemma 3, removing activation quantization specifically on decode improves accuracy on decode-heavy tasks without increasing inference cost. Training separate compute-native prefill weights accelerates prompt processing relative to weight-only inference while matching or exceeding its accuracy at 2-3-bit decode on both decode-heavy and prefill-heavy tasks. With released Qwen3.8-27B GGUF decoders, training an NVFP4 prefiller improves 1-bit accuracy by 32.5 points on MMLU-Pro and 35.3 on MMMU-Pro without modifying the decode checkpoint. To accommodate the additional checkpoint on a single device, offloaded disaggregated prefill (ODP) streams its weights from SSD, amortizing loading over prompt length. On the same 27B model, ODP delivers a 1.78x time-to-first-token speedup over the weight-only baseline at 8K prompt length in llama.cpp. We evaluate accuracy under disaggregated serving in vLLM and further validate shared-weight format disaggregation through post-training quantization on models up to 2.8T parameters.
Problem

Research questions and friction points this paper is trying to address.

disaggregated quantization
large language models
prefill and decode
quantization
inference cost
Innovation

Methods, ideas, or system contributions that make the work stand out.

disaggregated quantization
prefill and decode specialization
activation quantization removal
offloaded disaggregated prefill
🔎 Similar Papers
No similar papers found.