RPFQ-ViT: Rotated Phase-Frame Quantization for Extremely Low-Bit Weights in Vision Transformers

πŸ“… 2026-10-03
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the severe accuracy degradation of Vision Transformers under extremely low-bit quantization, caused by geometric mismatches between scalar codebooks and projection directions. To overcome this, we propose Rotational Phase Frame Quantization (RPFQ), which introduces a novel two-dimensional phase-plane channel pairing mechanism that maps paired channels onto the complex plane to preserve directional information. RPFQ further incorporates learnable rotation matrices, phase anchor learning, and residual refinement, recovering magnitudes via lightweight scaling. As a plug-and-play module, it requires no modifications to standard computational graphs. Combined with quantization-aware training, our method achieves 79.33% Top-1 accuracy on ImageNet under W2A4 settings, yielding 5.4–7.1Γ— model compression and 1.4–1.6Γ— latency reduction on mobile devices, thereby enabling high-accuracy, low-bit deployment.
πŸ“ Abstract
Vision Transformers (ViTs) achieve strong performance on image recognition and mobile vision applications, but their high-dimensional linear projections and attention computations still impose substantial storage and inference costs. Extremely low-bit quantization is a promising solution, yet ViTs often suffer severe accuracy degradation because conventional real-valued scalar codebooks are poorly matched to the directional geometry of Transformer projections. We present RPFQ-ViT, a Rotated Phase-Frame Quantization method that quantizes paired channels in two-dimensional phase planes, enabling low-bit codes to better preserve projection directions while recovering magnitude with lightweight scaling. RPFQ-ViT serves as a drop-in QAT replacement for nn.Linear and does not modify the standard real-valued attention, normalization, or activation computation graph. On ImageNet-1K, RPFQ-ViT-B/16 reaches 79.33% Top-1 / 94.48% Top-5 under W2/A4, Swin-T reaches 79.30% Top-1 / 94.79% Top-5 under W2/A8, and DeiT-S reaches 77.41% Top-1 / 93.11% Top-5 under W2/A8. Ablations, phase-geometry analysis, and direction-preservation metrics show that channel pairing, learnable rotation, phase-anchor learning, and residual phase refinement each improve quantization quality. We further deploy RPFQ-ViT image-classification models on native iOS and Android runtime stacks; with 2-bit packed weights, model size shrinks by roughly $5.4$-$7.1\times$ relative to FP32 and end-to-end on-device latency drops by $1.4$-$1.6\times$. All ImageNet results trained in our codebase use a matched 300-epoch recipe and are reported as mean accuracies over three independent runs. These results show that RPFQ-ViT provides a favorable trade-off among accuracy, compression, and practical mobile deployment for extremely low-bit ViTs.
Problem

Research questions and friction points this paper is trying to address.

Vision Transformers
Extremely low-bit quantization
Accuracy degradation
Storage and inference costs
Innovation

Methods, ideas, or system contributions that make the work stand out.

Rotated Phase-Frame Quantization
Vision Transformers
Extremely Low-Bit Quantization
Quantization-Aware Training
Mobile Deployment
πŸ’Ό Related Jobs
No related jobs found.
M
Mengyuan Fan
Peking University
B
Bokai Huang
Peking University
J
JiaMing Pan
Peking University
X
Xiaokun Yuan
Peking University
P
Peizhuang Cong
Peking University
Z
Zhewen Tan
Peking University
Tong Yang
Tong Yang
Peking University, Beijing, China. PKU. εŒ—δΊ¬ε€§ε­¦
SketchNetwork measurementBloom filterIP lookupHash Table