XFP: Quality-Targeted Adaptive Codebook Quantization with Sparse Outlier Separation for LLM Inference

📅 2026-05-14
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of balancing accuracy and efficiency in large language model inference under memory constraints, particularly for massive models with mixture-of-experts architectures, where conventional quantization methods rely on handcrafted bit-width assignments, calibration data, or Hessian information. The authors propose XFP, a dynamic quantization framework that introduces a novel channel-wise cosine similarity–based adaptive mechanism requiring neither calibration nor Hessian computation. XFP employs an H-Process to automatically search for optimal configurations satisfying both quality and memory constraints, decomposing weights into sparse fp16 outlier residuals and dense sub-byte indices, while co-designing codebooks, outlier handling, and fused decoding kernels. On Qwen3.5-122B, XFP achieves 138 tokens/s and 94.49% GSM8K accuracy—49% faster than Marlin INT4. On Qwen3.5-397B, it fits within 2×96GB GPUs at ~3.4 effective bits, delivering 100.9 tokens/s in long-context decoding, outperforming INT4 with expert pruning in both accuracy and efficiency.
📝 Abstract
We introduce XFP, a dynamic weight quantizer for LLM inference that inverts the conventional workflow: the operator specifies reconstruction quality floors on per-channel cosine similarity (one strict floor for attention and shared experts, one lazy floor for routed-expert MoE); XFP determines codebook size, outlier budget, and packing per layer automatically -- no Hessian, no calibration data, no manual bit-width selection. Each weight matrix is decomposed into a sparse fp16 outlier residual and a dense sub-byte index tensor into a per-group learned codebook. Two storage modes share one auto-select frontend and one fused decode kernel: V2 (per-channel Lloyd) and V2a (shared library of L=32 codebooks per layer). On Qwen3.5-122B-A10B under V2, XFP reaches 138 tok/s single-stream decode on workstation hardware (RTX PRO 6000 Blackwell, TP=2) at 94.49% GSM8K strict-match (3 seeds, n=3957), and is 49% faster than Marlin INT4 at TP=1. For models that do not fit in the target memory envelope, we present the H-Process: a quality-driven iteration over the two cosine thresholds that finds the operating point at which the model just fits while still producing sensible output. Three constraints define its search space: the operator-set thresholds, an OOM boundary at quantize-on-load, and a garbage boundary in generation (cosine similarity steers; benches verify). On Qwen3.5-397B-A17B (512 routed experts/layer), the H-Process fits the full expert population into 2x96 GB at ~3.4 effective bits and delivers 100.9 tok/s long-output decode at 66.72% GSM8K strict-match on the full 1319-problem set (single seed at submission; multi-seed evaluation in progress), exceeding INT4 with routed-expert pruning on memory, throughput, and accuracy simultaneously.
Problem

Research questions and friction points this paper is trying to address.

LLM inference
weight quantization
memory constraint
outlier handling
quality-targeted
Innovation

Methods, ideas, or system contributions that make the work stand out.

adaptive quantization
sparse outlier separation
quality-targeted inference
codebook learning
memory-constrained LLM
🔎 Similar Papers
2023-09-27International Conference on Learning RepresentationsCitations: 7
2024-06-17Neural Information Processing SystemsCitations: 14
💼 Related Jobs
No related jobs found.
T
Thomas Witt
Gemini Stiftung Leipzig