VisionMX: Unlocking Microscaling Post-Training Quantization for Vision Models

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the significant performance degradation incurred when applying Microscaling (MX) quantization to vision models, which stems from block-scale representation, weight misalignment, and insufficient utilization of activation signs. To overcome these challenges, this work systematically identifies the three error sources and proposes the first post-training quantization scheme specifically tailored for vision models under the MX format. The proposed approach integrates a bounded weight rounding algorithm with a collapsible affine correction technique, precisely compensating for quantization errors while preserving the highly efficient inference properties of MX representations. Extensive experiments demonstrate that this method substantially outperforms naive conversion and existing baselines across image classification and object detection tasks. Notably, it effectively recovers performance losses in sensitive architectures, achieving an optimal balance between model accuracy and inference efficiency.
📝 Abstract
Microscaling (MX) formats are emerging as a hardware-supported approach to efficient training and inference. They combine low-precision elements with shared block scales, but their impact on vision models remains underexplored. We systematically investigate post-training MX quantization across vision models and tasks. An analysis of direct conversion identifies three sources of error: block-scale representation, the poor alignment of some small convolutional weight tensors with nonuniform element grids, and the underuse of signed codes by nonnegative activations. These findings motivate VisionMX, a post-training MX quantization method that optimizes bounded weight rounding and applies a foldable affine correction to activations. We evaluate VisionMX across image classification, object detection, semantic segmentation, and low-light image enhancement using several MX-style formats. It improves on direct conversion and the evaluated post-training quantization baselines, with the largest performance recoveries in architectures most sensitive to MX conversion
Problem

Research questions and friction points this paper is trying to address.

Microscaling
Post-Training Quantization
Vision Models
MX Formats
Innovation

Methods, ideas, or system contributions that make the work stand out.

Microscaling Quantization
Post-Training Quantization
Vision Models
Bounded Weight Rounding
Foldable Affine Correction
🔎 Similar Papers
No similar papers found.