two-level activation scaling

Designs and implements a two-tier activation scaling mechanism that applies both tensor-level and block-level scale factors (a two-level scale hierarchy, e.g., for NVFP4-style 4-bit activations) to restore activation dynamic range and reduce quantization error across blocks. Builds and evaluates quantization pipelines and calibration/analysis tools that compute and apply these scales to enable accurate low-bit activations without retraining.

two-levelactivationscaling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.14
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Four Over Six: More Accurate NVFP4 Quantization with Adaptive Block Scaling

Dec 01, 2025
JC
Jack Cook
🏛️ Massachusetts Institute of Technology | NVIDIA

To address forward/backward divergence and degraded inference performance caused by NVFP4 low-precision quantization in large language model (LLM) training, this paper proposes the Four Over Six (4/6) quantization method. The core innovation lies in dynamically evaluating two candidate scaling factors per data block, prioritizing representation accuracy near the maximum absolute value region to substantially mitigate performance degradation dominated by FP4’s inherent quantization error. Built upon adaptive block-wise floating-point scaling, 4/6 is compatible with both forward and backward passes, as well as multiple post-training quantization paradigms, and features an optimized implementation for NVIDIA Blackwell architectures. Experiments demonstrate stable convergence during Transformer- and hybrid-architecture pretraining, with training loss closely matching BF16 baselines and consistent improvements in downstream task accuracy.

Addresses NVFP4 quantization errors causing training divergence and inference degradation.Enables efficient NVFP4 training on GPUs while enhancing model accuracy.Introduces adaptive block scaling to improve representation of near-maximal values.

This work addresses the significant gap between existing AbsMax-based block-wise scale initialization and the optimal solution in NVFP4 quantization, which severely limits the accuracy of 4-bit large language models. To overcome this limitation, we propose ScaleSweep, a method that efficiently searches for the optimal block scale within the first-ever derived theoretical upper and lower bounds on mean squared error (MSE) and weighted MSE (WMSE) for NVFP4. By minimizing reconstruction error, ScaleSweep enables high-accuracy post-training quantization with negligible computational overhead. Evaluations on Llama and Qwen models demonstrate that ScaleSweep substantially outperforms existing approaches, maintaining over 93% of the original model performance even when fully quantizing weights, activations, KV caches, and query states to 4 bits.

block scale initializationLLMsNVFP4

This work addresses the significant accuracy degradation caused by existing 4-bit quantization formats—such as NVFP4—when representing large-magnitude values. To mitigate this issue, the authors propose a block-wise adaptive mixed-precision quantization scheme that dynamically selects between INT4 and FP4 representations for every group of 16 values, reusing the sign bit of the shared scaling factor to indicate the format type. The approach is generalized to other bit widths, yielding a unified IFx family (e.g., IF3, IF6). Coupled with E4M3 scaling factors and a dedicated IF4 multiply-accumulate unit, the method enables efficient hardware deployment. Experiments demonstrate that IF4 consistently outperforms current 4-bit quantization strategies in both training and post-training settings, substantially reducing language modeling loss and improving accuracy across multiple downstream tasks.

4-bit formatblock-scaled data typeslarge language models

This work addresses the performance degradation of large-scale text-to-video diffusion Transformers—such as Wan2.1-T2V-14B—under uniform quantization, which arises from significant heterogeneity in activation distributions between boundary and intermediate layers. The study presents the first systematic analysis of layer-wise activation distribution differences in video DiTs and introduces a data-driven, boundary-preserving post-training quantization (PTQ) strategy: the middle 35 modules are quantized to W8A8 HiFloat8, while the first and last five modules retain BF16 precision. Evaluated on Ascend 910B, this approach achieves lossless compression, matching or slightly surpassing the BF16 baseline across all five VBench dimensions, thereby validating the efficacy of boundary protection. Furthermore, under this configuration, quantization-aware training (QAT) demonstrates no significant advantage over PTQ.

activation distributionboundary blocksquantization

QuEST: Low-bit Diffusion Model Quantization via Efficient Selective Finetuning

Feb 06, 2024
HW
Haoxuan Wang
🏛️ Illinois Institute of Technology | Houmo AI

Diffusion models face significant challenges in low-bit quantization deployment—including imbalanced activation distributions, temporal information distortion, and sensitivity to perturbations—making it difficult for existing methods to preserve generation quality under 4-bit weight/activation (W4A4) quantization. To address this, we propose a selective fine-tuning paradigm: leveraging module-level sensitivity analysis to identify critical temporal layers and sensitive components, and integrating adaptive activation distribution calibration with temporal information preservation. This enables stable training with minimal parameter updates. Our approach achieves the first end-to-end W4A4 quantization of Stable Diffusion capable of generating photorealistic images. It attains state-of-the-art performance across three high-resolution generation benchmarks, significantly improving quantization stability and visual fidelity. This work establishes an efficient and practical pathway for deploying diffusion models on resource-constrained edge devices.

Address high memory and computational overhead in diffusion modelsMitigate performance degradation in quantized layers selectivelyOvercome challenges in efficient low-bit quantization

Latest Papers

What's happening recently
View more

FP4 quantization struggles to preserve the accuracy of large language models due to its rigid coupling of quantization and dequantization scales and constraints imposed by hardware-supported discrete formats. This work proposes FOCUS, a novel framework that, for the first time, decouples these scale constraints by introducing Coupled-Relaxed Scaling (CRS) with learnable full-precision coefficients and sub-block-level Dual-Granularity Scaling (DGS). These innovations enhance model accuracy while maintaining compatibility with MXFP4/NVFP4 hardware formats. Leveraging post-training quantization, FOCUS achieves state-of-the-art FP4 performance across diverse large language models and benchmarks without incurring any additional inference overhead.

accuracy preservationFP4 quantizationhardware compatibility

This work presents the first application of MixQ in conjunction with SmoothQuant to enable efficient 4-bit (HiF4/MXFP4) inference for the Wan2.2-I2V-A14B image-to-video large model. To address heavy-tailed activation distributions, the authors propose a dual-branch quantization architecture that combines channel-wise smoothing to compress dynamic ranges with block-wise HiF4 packing and dual-branch GEMM. After calibration, outlier columns are retained in higher precision while the remaining channels undergo strict W4A4 quantization. Evaluated on VBench I2V, the method achieves performance within only 2–3.5% of FP16 across most metrics, substantially outperforming the native HiFloat4 baseline—which degrades by approximately 5%—and notably enhances motion smoothness in generated videos.

activation outlierslarge modellow-bit inference