HiLoRe: What to Store, Compress, or Recompute for Efficient GRPO Training
This study addresses the memory waste and recomputation overhead caused by fixed checkpointing strategies in GRPO training by proposing HiLoRe, an adaptive memory management method. For the first time, HiLoRe integrates the analytical update structure of GRPO into state fidelity allocation. By modeling policy update exposure and predicting approximation risk via loss coefficients, it dynamically allocates resources across high-precision storage, low-precision compression, and deterministic recomputation, thereby achieving memory optimization under risk-budget calibration. Experiments demonstrate that under memory-constrained conditions, HiLoRe improves Actor update throughput by 13.5% compared to gradient checkpointing (GC), while maintaining downstream task performance within a 0.6 percentage point margin. These results confirm that the proposed approach effectively balances training efficiency with model quality.