QuantForge: Discovering Residual Decompositions for MXFP4 Post-Training Quantization

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges in MXFP4 quantization, where coordinate transformation is coupled with residual propagation and single performance metrics inadequately guide search. We propose an LLM-driven program evolution framework featuring a novel residual compilation mechanism that translates discriminative evidence from competing interpretations into targeted code modifications. This approach eliminates erroneous paths while preserving valid logic, enabling efficient automated search for W4A4 post-training quantization algorithms. The method discovers the HiRes quantizer, which achieves the lowest robust fitting error (0.093) across seven tasks and consistently meets transfer objectives in all six runs under equivalent budgets, significantly outperforming existing baselines.
📝 Abstract
Four-bit post-training quantization can reduce the memory demands of large language models, but preserving accuracy under strict MXFP4 W4A4 requires coordinating several design choices. Coordinate transforms change block-encoding errors, which in turn affect the residuals propagated through the network. The useful algorithmic decomposition is therefore not fully known before search. LLM-driven program evolution offers a way to explore these choices, but performance scores alone do not explain which design should change next. We introduce QuantForge, a PTQ discovery system that records competing explanations, selects controls that distinguish them, and checks that successor code implements the resulting conclusions. This residual compilation guides program revisions while retaining useful programs even when their original explanations are rejected. Remeasuring the revised program reveals the next error to address. This process discovers HiRes, a fixed MXFP4 quantizer that shapes coordinates, refines legal code assignments, and recovers errors along attention and MLP paths. Each stage acts on residuals measured after the preceding stage has executed. Across seven tasks, HiRes achieves the lowest seven-model Robust Fit (0.09300) and the lowest quantized Fit-7 at 32B. In matched-budget comparisons of LLM-driven program evolution, each with 240 evaluator calls, QuantForge reaches a held-out transfer target in six of eight runs, compared with three each for textual memory and reflection memory, and one for score-only evolution, despite evaluating fewer new programs. These results show that QuantForge improves the discovery of transferable PTQ algorithms by turning controlled evidence into subsequent program changes.
Problem

Research questions and friction points this paper is trying to address.

Post-Training Quantization
MXFP4
Large Language Models
Program Evolution
Residual Decomposition
Innovation

Methods, ideas, or system contributions that make the work stand out.

Post-Training Quantization
MXFP4
LLM-driven Program Evolution
Residual Compilation
HiRes
💼 Related Jobs
No related jobs found.
Q
Qiulin Shang
Peking University
Z
Zhoutong Wu
Peking University
J
Jie Hu
Peking University
Kun Yuan
Kun Yuan
Center for Machine Learning Research, Peking University
distributed signal processinglarge-scale optimizationmachine learning