Score
Design and implement scoring kernels and data structures that compute similarities between queries and product-quantized database vectors by performing fused lookup-table operations on quantized codes, including quantized lookup and shared-memory codebook lookups to reduce external memory I/O. Analyze and optimize these fused PQ lookup pipelines to minimize bandwidth and latency (e.g., cutting high-bandwidth memory I/O) while preserving top‑k / max‑similarity ranking fidelity.
Large language model (LLM) inference is bottlenecked by GPU memory bandwidth, particularly under non-uniform low-bit (e.g., 3-bit) lookup table (LUT) quantization, where fused dequantization and matrix multiplication suffers from poor computational efficiency. To address this, we propose FLUTE, a novel inference engine featuring the first CUDA kernel supporting non-divisible-bit-width LUT quantization. FLUTE reduces bit-manipulation overhead via offline weight reconstruction and alleviates shared-memory bandwidth pressure through LUT vectorization and redundant loading. It natively supports weight-only quantization, NormalFloat extensions, and bandwidth-aware scheduling. Experiments show that, under batch size < 32 and group size = 128, FLUTE’s kernel achieves 2–4× speedup over state-of-the-art GEMM kernels. When integrated into the LLaMA-3 quantization backend, FLUTE delivers 1.5–2× end-to-end throughput improvement while preserving accuracy competitive with strong baselines.
This work addresses the high memory and computational costs associated with approximate nearest neighbor (ANN) search in large-scale, high-dimensional datasets. To overcome the limitations of single-machine resources, the authors propose a divide-and-conquer parallelization framework built on Dask that efficiently integrates product quantization (PQ) with inverted indexing. By distributing computation and storage across multiple nodes, the method significantly reduces resource demands while preserving search accuracy. As a result, the computational overhead of large-scale high-dimensional ANN search is brought down to levels comparable to those of medium-scale datasets, enabling scalable and efficient approximate retrieval.
This work addresses the limitation of existing KV cache quantization methods, which, despite reducing storage, still require dequantization to high-precision floating-point numbers and thus fail to alleviate the memory bandwidth bottleneck in attention computation. To overcome this, the paper introduces— for the first time—product quantization and asymmetric distance computation from vector retrieval into the Transformer attention mechanism. By leveraging subspace decomposition, codebook learning, and lookup-table-based operations, the approach transforms memory-bound attention into compute-bound without modifying the model architecture or requiring retraining. Evaluated on GPT-2, the method achieves 95.7% output fidelity at a 64× compression ratio (95.0% at 32×) and maintains a Spearman rank correlation coefficient ρ > 0.95 in attention scores. Both theoretical analysis and experiments confirm its effectiveness for sequence lengths up to 1024.
To address the limitations of fixed-bit quantization, accuracy degradation, and high query latency in Approximate k-Nearest Neighbor (AKNN) search within high-dimensional Euclidean spaces, this paper proposes Multi-Granularity Residual Quantization (MRQ). MRQ decouples the number of quantization bits from vector dimensionality for the first time, enhances distance correction accuracy via data distribution modeling, and integrates adaptive vector quantization, data-driven distance correction, efficient quantized distance computation, and error-bound optimization. Compared with state-of-the-art graph-based and quantization-based methods (e.g., RaBitQ), MRQ achieves a threefold speedup in query latency while maintaining identical retrieval accuracy and reducing code length to one-third. This significantly improves index configurability and practical applicability for large-scale AKNN search.
To address the GPU memory bottleneck induced by KV cache in long-context LLM inference, this work pioneers the application of Product Quantization (PQ) to KV cache compression, framing it as an approximate nearest neighbor search in embedding space. We propose an overlapping block partitioning scheme coupled with a hierarchical caching mechanism that eliminates extraneous computation and communication overhead across both prefill and decode stages. Our method preserves model quality while substantially reducing service latency: it achieves a 4.60% score improvement on InfiniteBench, outperforms state-of-the-art methods in both prefill and decode latency, and enables efficient inference over context lengths exceeding 10,000 tokens. The core contributions are (i) a PQ-driven KV compression paradigm that drastically reduces memory footprint without accuracy degradation, and (ii) a zero-overhead hierarchical scheduling design that seamlessly integrates compression into the inference pipeline.
This study addresses the substantial storage and access overhead of long-context KV caches, noting that existing rotation-based quantization methods reduce bit-width but disperse query energy, thereby hindering efficient channel pruning. To overcome this limitation, this work proposes the Dual-QK framework, which leverages paired non-orthogonal transformations to balance key quantization scales and concentrate query energy. By integrating calibration statistics, partial key whitening, and a Channel-0 protection mechanism, Dual-QK achieves synergistic optimization of low-bit quantization and dynamic channel pruning. Experimental results demonstrate that the proposed method outperforms OSCAR in accuracy at 40% sparsity. Furthermore, on 128K context lengths, it attains 6.8× compression and an 8.3× reduction in memory reads, yielding up to a 3.75× improvement in decoding throughput.
This study addresses the prohibitive storage overhead of large look-up tables (LUTs) in edge devices by proposing CompressedLUT, a lossless compression scheme coupled with an efficient hardware decoding architecture. Methodologically, the approach integrates matrix factorization, self-similarity mining, and multi-level compression techniques to maximize the compression ratio with zero precision loss. Furthermore, a lightweight decoder based on addition, arithmetic right shifts, and micro-LUTs is designed to achieve low area consumption and high throughput. Experimental results demonstrate that this scheme significantly reduces hardware resource utilization in scenarios such as nonlinear function computation on FPGAs. The associated tools have been released as open source.
This study addresses the trade-off in tabular in-context learning (ICL) between the accuracy degradation of fixed subsets and the low throughput of dynamic retrieval, proposing the QCOC method. QCOC introduces a novel commutativity-based operator compression mechanism that compiles contextual KV caches into shared compact prototypes. By integrating joint KV clustering with attention vector anchoring calibration and optimizing prototype values via closed-form solutions, it achieves efficient reuse while preserving accuracy. Experiments on OpenML datasets demonstrate that QCOC attains leading accuracy, yielding a 10.5× cache compression ratio and a 508× inference speedup over dynamic retrieval, effectively overcoming the accuracy-efficiency bottleneck.
This study addresses the memory bandwidth bottleneck in long-context inference caused by the mismatch between KV cache storage precision and decoding query requirements. We propose ReadKV, a method that stores KV caches via progressive encoding and introduces a pioneering query-adaptive quantization mechanism. This mechanism dynamically allocates reading budgets based on the current query to reconstruct keys and values on demand, achieving precise retrieval through calibrated distortion objective optimization and a dynamic prefix allocation algorithm. Both theoretical analysis and experiments demonstrate that, under a fixed budget, query-dependent access strategies strictly outperform query-agnostic approaches. Compared to full 4-bit reading, ReadKV achieves superior accuracy with only a 0.66% increase in C4 perplexity, while reducing logical reads by 75% and latency by 39%.
This study addresses the KV cache memory bottleneck in long-context LLM inference, where existing quantization methods suffer significant performance degradation under extreme 1-bit compression. To overcome this, we propose TaSQ, an efficient 1-bit KV cache compression method leveraging a tailored vector quantization target space. Its core innovations include query-guided weighting, cross-head normalization, and covariance-aware grouping to precisely model activation statistics, alongside RoPE-compatible transformations and adaptive channel grouping to preserve representational capacity under extreme compression. Implemented within the SGLang framework, TaSQ outperforms existing low-bit baselines across multiple benchmarks, enabling a 14× larger batch size and achieving a 1.87× higher peak throughput than BF16 while maintaining stable inference.