TRACE: Rollout-Guided Quantization-Aware Training for FP4 Reinforcement Learning of MoE Language Models

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the significant quantization discrepancy between training and inference pathways, as well as the high computational overhead, in FP4 reinforcement learning. To this end, we propose TRACE, a novel framework whose core innovation lies in an inference-pathway-guided quantization-aware training mechanism that directly aligns quantization decisions across both pathways. Furthermore, a selective caching scheme is designed to substantially reduce storage and communication costs. Experimental results demonstrate that the proposed method achieves a 5.4× inference speedup across multiple large-scale Mixture-of-Experts (MoE) models while delivering performance comparable to BF16 baselines. This work establishes a new paradigm for efficient low-bit MoE reinforcement learning.
📝 Abstract
Reinforcement learning (RL) for post-training large language models (LLMs) incurs substantial computation and memory overhead during rollout generation, which motivates low-precision rollout for efficient RL training. However, existing FP4 RL methods suffer from a key limitation: they primarily optimize quantization accuracy on the training and rollout paths independently rather than directly reducing the discrepancy between the two quantized execution paths. In this work, we propose TRACE (Train-Rollout Quantization Alignment via Compact GuidancE), an FP4 quantization framework for RL training of Mixture-of-Experts (MoE) language models that addresses the limitation of existing FP4 RL methods. TRACE incorporates rollout-guided quantization-aware training that uses rollout-side quantization outcomes to guide training-side FP4 rounding decisions, directly reducing train-rollout discrepancy. Moreover, TRACE adopts an efficient quantization-information caching scheme that selectively retains mantissa and scale information from deeper layers to reduce the storage and communication overhead introduced by rollout guidance. We evaluate TRACE on four large-scale MoE language models across reasoning, coding, and long-horizon RL tasks. Our results demonstrate that TRACE enables joint FP4 weight/activation and FP4 KV-cache rollout with RL performance comparable to BF16 rollout, while achieving up to 5.4xrollout speedup and strong final FP4 performance compared with post-hoc FP4 quantization of BF16-trained policies.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
FP4 Quantization
Mixture-of-Experts
Train-Rollout Discrepancy
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

FP4 Quantization
Reinforcement Learning
Mixture-of-Experts
Quantization-Aware Training
Rollout-Guided Alignment
🔎 Similar Papers
2024-08-10AAAI Conference on Artificial IntelligenceCitations: 30
X
Xin Wang
Alibaba Token Hub, Alibaba Group
H
Hao Yu
Alibaba Token Hub, Alibaba Group
Z
Zhengyang Zhuge
Alibaba Token Hub, Alibaba Group
B
Bochao Mao
Alibaba Token Hub, Alibaba Group
Z
Zheng Li
Alibaba Token Hub, Alibaba Group
Junda Feng
Junda Feng
Unknown affiliation
Y
Yuyan Luo
Alibaba Token Hub, Alibaba Group
Y
Yi Zhang
Alibaba Token Hub, Alibaba Group
Y
Yizhong Cao
Alibaba Token Hub, Alibaba Group
Mi Zhang
Mi Zhang
Associate Professor, Director of AIoT and MLSys Lab, The Ohio State University
Edge AIAIoTMachine Learning SystemsGenerative AIMobile Computing
D
Dayiheng Liu
Alibaba Token Hub, Alibaba Group
J
Jianwei Zhang
Alibaba Token Hub, Alibaba Group