TRIAGE: Direction-Aware Mismatch Stabilization of Native NVFP4 Reinforcement Learning
This study addresses the training instability arising from learner-sampler misalignment in low-precision policy optimization. We propose a direction-aware stabilization method that employs a piecewise diagnostic mechanism to reveal the interaction between such mismatches and gradient directions, alongside selective rebalancing and bounded repair algorithms to correct severe deviations. This approach enables native NVFP4 (W4A4) quantized training, achieving stable optimization while preserving forward inference efficiency. Experimental results demonstrate that the proposed method attains full-precision performance on mathematical reasoning benchmarks, yielding throughput improvements of up to 2.3× over the BF16 baseline.