AlignQuant: Tile-Aligned Mixed-Precision Quantization for Efficient LLM Generation

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge in fine-grained mixed-precision quantization where mismatches between local precision selection and GPU memory-compute units hinder the translation of model compression into practical speedups. To this end, we propose AlignQuant, a framework that introduces a novel shared partitioning mechanism based on GPU-compatible two-dimensional weight blocks, unifying precision allocation, storage, and execution at the block level. By integrating joint prefill-decode calibration with phase-normalized scoring optimization, AlignQuant reconciles flexible precision assignment with efficient execution on commodity GPUs. Extensive experiments on large language models ranging from 3B to 14B parameters demonstrate that our method achieves up to 2.5× generation speedup without compromising model quality. Furthermore, its effectiveness is validated across multiple GPU architectures and in 64K long-context scenarios.
📝 Abstract
Fine-grained mixed-precision quantization promises efficient large language model inference, but local precision choices can conflict with regular GPU storage and computation units. This precision-boundary mismatch limits the translation of compression into practical acceleration. We introduce AlignQuant, a post-training quantization method that uses GPU-compatible two-dimensional weight tiles as the common unit of precision allocation, compact storage, and execution. This shared partition lets precision follow sensitivity within output channels. Joint prefill/decode calibration scores precision reductions using projection-output perturbations weighted by language-model loss gradients under quantized activations. Phase-normalized scores prioritize higher precision for tiles important to either phase under a model-wide weight-storage budget. Each tile stores one selected representation, while phase-specialized kernels reuse the packed model and expand lower-bit weights for INT8 computation with 8-bit activations. Across four LLMs spanning 3B to 14B parameters, AlignQuant achieves up to $2.50\times$ generation speedup over BF16 while preserving model quality. Evaluations further cover three GPUs and contexts up to 64K tokens. These results show that local precision flexibility and regular GPU execution can coexist through a shared tile unit. The implementation is available at https://github.com/HanzhiZhang-Ulrica/AlignQuant.
Problem

Research questions and friction points this paper is trying to address.

mixed-precision quantization
large language models
GPU efficiency
precision-boundary mismatch
inference acceleration
Innovation

Methods, ideas, or system contributions that make the work stand out.

Mixed-Precision Quantization
Tile-Aligned Partitioning
Post-Training Quantization
LLM Inference Acceleration
Joint Calibration
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.