π€ AI Summary
This study addresses the severe accuracy degradation of Vision Transformers under extremely low-bit quantization, caused by geometric mismatches between scalar codebooks and projection directions. To overcome this, we propose Rotational Phase Frame Quantization (RPFQ), which introduces a novel two-dimensional phase-plane channel pairing mechanism that maps paired channels onto the complex plane to preserve directional information. RPFQ further incorporates learnable rotation matrices, phase anchor learning, and residual refinement, recovering magnitudes via lightweight scaling. As a plug-and-play module, it requires no modifications to standard computational graphs. Combined with quantization-aware training, our method achieves 79.33% Top-1 accuracy on ImageNet under W2A4 settings, yielding 5.4β7.1Γ model compression and 1.4β1.6Γ latency reduction on mobile devices, thereby enabling high-accuracy, low-bit deployment.
π Abstract
Vision Transformers (ViTs) achieve strong performance on image recognition and mobile vision applications, but their high-dimensional linear projections and attention computations still impose substantial storage and inference costs. Extremely low-bit quantization is a promising solution, yet ViTs often suffer severe accuracy degradation because conventional real-valued scalar codebooks are poorly matched to the directional geometry of Transformer projections. We present RPFQ-ViT, a Rotated Phase-Frame Quantization method that quantizes paired channels in two-dimensional phase planes, enabling low-bit codes to better preserve projection directions while recovering magnitude with lightweight scaling. RPFQ-ViT serves as a drop-in QAT replacement for nn.Linear and does not modify the standard real-valued attention, normalization, or activation computation graph. On ImageNet-1K, RPFQ-ViT-B/16 reaches 79.33% Top-1 / 94.48% Top-5 under W2/A4, Swin-T reaches 79.30% Top-1 / 94.79% Top-5 under W2/A8, and DeiT-S reaches 77.41% Top-1 / 93.11% Top-5 under W2/A8. Ablations, phase-geometry analysis, and direction-preservation metrics show that channel pairing, learnable rotation, phase-anchor learning, and residual phase refinement each improve quantization quality. We further deploy RPFQ-ViT image-classification models on native iOS and Android runtime stacks; with 2-bit packed weights, model size shrinks by roughly $5.4$-$7.1\times$ relative to FP32 and end-to-end on-device latency drops by $1.4$-$1.6\times$. All ImageNet results trained in our codebase use a matched 300-epoch recipe and are reported as mean accuracies over three independent runs. These results show that RPFQ-ViT provides a favorable trade-off among accuracy, compression, and practical mobile deployment for extremely low-bit ViTs.