🤖 AI Summary
This study addresses the memory waste and recomputation overhead caused by fixed checkpointing strategies in GRPO training by proposing HiLoRe, an adaptive memory management method. For the first time, HiLoRe integrates the analytical update structure of GRPO into state fidelity allocation. By modeling policy update exposure and predicting approximation risk via loss coefficients, it dynamically allocates resources across high-precision storage, low-precision compression, and deterministic recomputation, thereby achieving memory optimization under risk-budget calibration. Experiments demonstrate that under memory-constrained conditions, HiLoRe improves Actor update throughput by 13.5% compared to gradient checkpointing (GC), while maintaining downstream task performance within a 0.6 percentage point margin. These results confirm that the proposed approach effectively balances training efficiency with model quality.
📝 Abstract
Group-relative policy optimization (GRPO) makes learner-side activations a major memory-computation bottleneck: gradient checkpointing reduces activation memory through recomputation, but fixed schedules can leave roughly 18 GB unused on a 48-GB GPU despite substantial recomputation overhead. Existing activation-management methods set state fidelity from execution cost, tensor properties, or generic compression sensitivity, without explicitly incorporating GRPO's analytic update structure into state-fidelity allocation. We formalize this dependence as policy-update exposure, linking the current GRPO loss coefficients to state-level approximation sensitivity. These coefficients are available before backward without an additional backward pass. We introduce HiLoRe, which allocates graph-attributed recovery units among high-precision storage, low-precision compression, and deterministic recomputation using measured recovery utility and update-conditioned approximation risk. It combines high-precision storage and deterministic recomputation with low-precision recovery under a calibrated risk budget. Across five model-task settings with 2K responses and memory<1.10 times GC's per-GPU actor-update peak, HiLoRe's actor-update throughput gains reach 13.5% over GC and 7.9% over the fastest evaluated baseline, with paired mean downstream-score differences below 0.6 percentage points.