Finer is Better (with the Right Scaling)

๐Ÿ“… 2026-05-08
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the paradoxical performance degradation of large language models under ultra-low-precision quantization (e.g., FP4) when using finer block sizes, revealing that the root cause lies in the mismatch between tensor value distributions and quantization intervalsโ€”not the block size itself. To resolve this, the authors propose a micro-scaling technique, a 4-over-6 quantization strategy, and a safeguard mechanism against scale-factor underflow, complemented by an exhaustive search to identify the optimal baseline configuration. Theoretical analysis demonstrates that quantization mean squared error strictly decreases with smaller block sizes. Extensive experiments show that the proposed method significantly improves downstream task perplexity across multiple mainstream large language models and achieves performance comparable to custom wide-exponent formats while remaining compatible with standard hardware-friendly representations.
๐Ÿ“ Abstract
Microscaling is a critical technique for preserving the quality of Large Language Models (LLMs) quantized to ultra-low precision formats. Intuitively, finer block sizes should yield lower quantization error; however, a paradox recently identified in the literature demonstrates that standard abs-max scaling can actually degrade model quality as block sizes shrink. In this work, we investigate the underlying mechanics of this phenomenon. We demonstrate that this degradation is not an inherent limitation of finer granularity, but is primarily driven by heavy-tailed tensor distributions interacting poorly with the coarse upper quantization bins of the FP4 element format. Specifically, we show that i) preventing the scaling factor from underflowing to zero mitigates localized errors, ii) targeted algorithmic interventions like the 4-over-6 methodology effectively correct the quantization geometry for large elements, and iii) a brute-force search establishes an optimal baseline, confirming that the theoretical Mean Squared Error (MSE) strictly improves with finer block sizes. Ultimately, our findings reveal a valuable interchangeability: applying the correct algorithmic recipe allows standard, hardware-compliant formats (like OCP E4M3) to match the performance of custom, wider-exponent formats (like UE5M3). We validate these results across several large language models, fully resolving the block size paradox and achieving robust downstream perplexity improvements.
Problem

Research questions and friction points this paper is trying to address.

quantization
block size paradox
large language models
ultra-low precision
microscaling
Innovation

Methods, ideas, or system contributions that make the work stand out.

microscaling
block size paradox
ultra-low precision quantization
heavy-tailed distributions
FP4 quantization
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
C
Clemens Schaefer
Google LLC, Mountain View, Ca
Gil Tabak
Gil Tabak
PhD Student in Applied Physics, Stanford University