RWKVQuant: Quantizing the RWKV Family with Proxy Guided Hybrid of Scalar and Vector Quantization

📅 2025-05-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the significant accuracy degradation in post-training quantization (PTQ) of RWKV models—caused by nonlinear operators impeding parameter fusion and uniform weight distributions undermining clustering efficacy—this paper proposes the first RWKV-specific PTQ framework. We introduce a novel coarse-to-fine weight proxy mechanism that adaptively coordinates scalar quantization (SQ) and vector quantization (VQ). Furthermore, to handle RWKV’s core element-wise multiplication operations, we design a codebook optimization algorithm that jointly models weight uniformity and outlier characteristics. Evaluated on RWKV-6-14B, our method achieves ~3-bit quantization with <1% accuracy loss and 2.14× inference speedup—marking the first breakthrough in balancing accuracy and efficiency for RNN-based model PTQ.

Technology Category

Machine Learning: Calibration & Uncertainty QuantificationSearch and Optimization: Learning to SearchNatural Language Processing: Learning & Optimization for NLP

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsWeb Mining and Content Analysis: Large pretrained models with web data
📝 Abstract
RWKV is a modern RNN architecture with comparable performance to Transformer, but still faces challenges when deployed to resource-constrained devices. Post Training Quantization (PTQ), which is a an essential technique to reduce model size and inference latency, has been widely used in Transformer models. However, it suffers significant degradation of performance when applied to RWKV. This paper investigates and identifies two key constraints inherent in the properties of RWKV: (1) Non-linear operators hinder the parameter-fusion of both smooth- and rotation-based quantization, introducing extra computation overhead. (2) The larger amount of uniformly distributed weights poses challenges for cluster-based quantization, leading to reduced accuracy. To this end, we propose RWKVQuant, a PTQ framework tailored for RWKV models, consisting of two novel techniques: (1) a coarse-to-fine proxy capable of adaptively selecting different quantization approaches by assessing the uniformity and identifying outliers in the weights, and (2) a codebook optimization algorithm that enhances the performance of cluster-based quantization methods for element-wise multiplication in RWKV. Experiments show that RWKVQuant can quantize RWKV-6-14B into about 3-bit with less than 1% accuracy loss and 2.14x speed up.
Problem

Research questions and friction points this paper is trying to address.

Quantization challenges in RWKV models due to non-linear operators
Uniform weight distribution complicates cluster-based quantization accuracy
Need for efficient PTQ framework to reduce RWKV model size
Innovation

Methods, ideas, or system contributions that make the work stand out.

Coarse-to-fine proxy for adaptive quantization selection
Codebook optimization for cluster-based quantization
3-bit quantization with minimal accuracy loss
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
C
Chen Xu
Houmo AI
Y
Yuxuan Yue
Houmo AI, Harbin Institute of Technology (Shenzhen)
Z
Zukang Xu
Houmo AI
X
Xing Hu
Houmo AI
Jiangyong Yu
Jiangyong Yu
houmo.ai
Z
Zhixuan Chen
Houmo AI
Sifan Zhou
Sifan Zhou
Southeast University
RoboticsM/LLMsSpatial AIQuantization
Zhihang Yuan
Zhihang Yuan
Bytedance
Efficient AIModel CompressionMLLM
D
Dawei Yang
Houmo AI