Score
Designs and implements a two-tier activation scaling mechanism that applies both tensor-level and block-level scale factors (a two-level scale hierarchy, e.g., for NVFP4-style 4-bit activations) to restore activation dynamic range and reduce quantization error across blocks. Builds and evaluates quantization pipelines and calibration/analysis tools that compute and apply these scales to enable accurate low-bit activations without retraining.
To address forward/backward divergence and degraded inference performance caused by NVFP4 low-precision quantization in large language model (LLM) training, this paper proposes the Four Over Six (4/6) quantization method. The core innovation lies in dynamically evaluating two candidate scaling factors per data block, prioritizing representation accuracy near the maximum absolute value region to substantially mitigate performance degradation dominated by FP4’s inherent quantization error. Built upon adaptive block-wise floating-point scaling, 4/6 is compatible with both forward and backward passes, as well as multiple post-training quantization paradigms, and features an optimized implementation for NVIDIA Blackwell architectures. Experiments demonstrate stable convergence during Transformer- and hybrid-architecture pretraining, with training loss closely matching BF16 baselines and consistent improvements in downstream task accuracy.
This work addresses the significant gap between existing AbsMax-based block-wise scale initialization and the optimal solution in NVFP4 quantization, which severely limits the accuracy of 4-bit large language models. To overcome this limitation, we propose ScaleSweep, a method that efficiently searches for the optimal block scale within the first-ever derived theoretical upper and lower bounds on mean squared error (MSE) and weighted MSE (WMSE) for NVFP4. By minimizing reconstruction error, ScaleSweep enables high-accuracy post-training quantization with negligible computational overhead. Evaluations on Llama and Qwen models demonstrate that ScaleSweep substantially outperforms existing approaches, maintaining over 93% of the original model performance even when fully quantizing weights, activations, KV caches, and query states to 4 bits.
This work addresses the significant accuracy degradation caused by existing 4-bit quantization formats—such as NVFP4—when representing large-magnitude values. To mitigate this issue, the authors propose a block-wise adaptive mixed-precision quantization scheme that dynamically selects between INT4 and FP4 representations for every group of 16 values, reusing the sign bit of the shared scaling factor to indicate the format type. The approach is generalized to other bit widths, yielding a unified IFx family (e.g., IF3, IF6). Coupled with E4M3 scaling factors and a dedicated IF4 multiply-accumulate unit, the method enables efficient hardware deployment. Experiments demonstrate that IF4 consistently outperforms current 4-bit quantization strategies in both training and post-training settings, substantially reducing language modeling loss and improving accuracy across multiple downstream tasks.
This work addresses the performance degradation of large-scale text-to-video diffusion Transformers—such as Wan2.1-T2V-14B—under uniform quantization, which arises from significant heterogeneity in activation distributions between boundary and intermediate layers. The study presents the first systematic analysis of layer-wise activation distribution differences in video DiTs and introduces a data-driven, boundary-preserving post-training quantization (PTQ) strategy: the middle 35 modules are quantized to W8A8 HiFloat8, while the first and last five modules retain BF16 precision. Evaluated on Ascend 910B, this approach achieves lossless compression, matching or slightly surpassing the BF16 baseline across all five VBench dimensions, thereby validating the efficacy of boundary protection. Furthermore, under this configuration, quantization-aware training (QAT) demonstrates no significant advantage over PTQ.
Diffusion models face significant challenges in low-bit quantization deployment—including imbalanced activation distributions, temporal information distortion, and sensitivity to perturbations—making it difficult for existing methods to preserve generation quality under 4-bit weight/activation (W4A4) quantization. To address this, we propose a selective fine-tuning paradigm: leveraging module-level sensitivity analysis to identify critical temporal layers and sensitive components, and integrating adaptive activation distribution calibration with temporal information preservation. This enables stable training with minimal parameter updates. Our approach achieves the first end-to-end W4A4 quantization of Stable Diffusion capable of generating photorealistic images. It attains state-of-the-art performance across three high-resolution generation benchmarks, significantly improving quantization stability and visual fidelity. This work establishes an efficient and practical pathway for deploying diffusion models on resource-constrained edge devices.
该研究针对现有量化方法依赖标量敏感度代理导致的精度损失问题,提出了一种基于激活意识和跨层优化的新量化方法CASA。
为提高大型语言模型推理效率,提出HBQ方法,通过层次化块量化和硬件高效设计解决精度与效率之间的权衡问题。
FP4 quantization struggles to preserve the accuracy of large language models due to its rigid coupling of quantization and dequantization scales and constraints imposed by hardware-supported discrete formats. This work proposes FOCUS, a novel framework that, for the first time, decouples these scale constraints by introducing Coupled-Relaxed Scaling (CRS) with learnable full-precision coefficients and sub-block-level Dual-Granularity Scaling (DGS). These innovations enhance model accuracy while maintaining compatibility with MXFP4/NVFP4 hardware formats. Leveraging post-training quantization, FOCUS achieves state-of-the-art FP4 performance across diverse large language models and benchmarks without incurring any additional inference overhead.
This work presents the first application of MixQ in conjunction with SmoothQuant to enable efficient 4-bit (HiF4/MXFP4) inference for the Wan2.2-I2V-A14B image-to-video large model. To address heavy-tailed activation distributions, the authors propose a dual-branch quantization architecture that combines channel-wise smoothing to compress dynamic ranges with block-wise HiF4 packing and dual-branch GEMM. After calibration, outlier columns are retained in higher precision while the remaining channels undergo strict W4A4 quantization. Evaluated on VBench I2V, the method achieves performance within only 2–3.5% of FP16 across most metrics, substantially outperforming the native HiFloat4 baseline—which degrades by approximately 5%—and notably enhances motion smoothness in generated videos.
本文针对4位浮点数预训练不稳定的问题,提出了一种结合E2M1载荷与无符号E5M3块尺度的方法,并通过选择性随机舍入和全FP4内部线性运算来实现更稳定的语言模型预训练。