OSFP4: Joint Optimization of Diagonal Smoothing and Block Scales for NVFP4 Quantization

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the severe accuracy degradation incurred by NVFP4 quantization during large language model (LLM) inference by proposing OSFP4. To the best of our knowledge, this method is the first to jointly optimize diagonal smoothing matrices and block scaling factors to minimize quantization error. Furthermore, it introduces a multiplicative dithered FP4 quantizer to establish a rigorous theoretical analysis framework, and incorporates GPTQ-based interference cancellation to further enhance performance. Experimental results demonstrate that OSFP4 significantly outperforms existing methods in average accuracy while retaining 94%–97% of the prefill throughput achieved by NVFP4, thereby realizing an effective balance between high precision and computational efficiency.
📝 Abstract
NVFP4 is an attractive datatype for large language model (LLM) inference, offering compact storage and native tensor-core acceleration. However, preserving accuracy using NVFP4 requires careful quantization. In this work we develop a novel quantization scheme called Optimized Smoothing and Scaling for NVFP4 (OSFP4). For each linear projection it uses a diagonal smoothing matrix whose entries are optimized to minimize the squared matrix-product quantization error under NVFP4, taking into account the rounding procedure that is used (either round-to-nearest, or GPTQ-style successive interference cancellation). This requires performing joint optimization on the smoothing entries as well as the block scales, which is facilitated by analyzing a multiplicative-dither FP4 quantizer instead of the fixed deterministic one. Experiments show that OSFP4 achieves the highest average accuracy among the evaluated competitors in the corresponding quantization settings, while retaining approximately 94-97\% of vendor NVFP4 prefill throughput on the measured workloads. Our code is available in https://github.com/neriahbd/OSFP4
Problem

Research questions and friction points this paper is trying to address.

NVFP4
quantization
large language models
inference accuracy
Innovation

Methods, ideas, or system contributions that make the work stand out.

NVFP4 Quantization
Diagonal Smoothing
Joint Optimization
Block Scales
Multiplicative Dither
N
Neriah Ben David
School of Computer Science and Engineering, Hebrew University of Jerusalem, Jerusalem, Israel
O
Ori Meir
School of Computer Science and Engineering, Hebrew University of Jerusalem, Jerusalem, Israel
Or Ordentlich
Or Ordentlich
Hebrew University of Jerusalem