LoGRA: Scaling LLM Reinforcement Learning with Low-Rank Gradient Sketches

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive memory overhead of the Adam optimizer during reinforcement learning (RL) post-training of large language models, which impedes scalable deployment. To overcome this limitation, we propose a memory-efficient training framework that synergistically integrates low-rank gradient sketching to compress gradients while preserving critical learning signals with a predictive KL step-size control mechanism to constrain policy update magnitudes. Experimental results demonstrate that the proposed approach reduces average training memory consumption by 45.7% without compromising model performance. Notably, it enables stable training for over 1,100 steps on a 27B-parameter model using a single eight-GPU node. By effectively circumventing the memory bottleneck inherent in dense Adam optimization, this work substantially enhances the scalability of RL-based post-training for large language models.
📝 Abstract
Reinforcement learning (RL) has greatly advanced the capabilities of large language models (LLMs), but its memory demands remain a barrier to broader adoption. We introduce LoGRA, an approach to RL post-training that reduces memory by retaining useful learning signals in low-rank gradient sketches. These compact representations support both model updates and efficient policy synchronization. To prevent overly large updates from disrupting learning, we complement gradient compression with predicted-KL step control, which estimates policy changes before applying each update and adjusts its magnitude accordingly. Across reasoning tasks, LoGRA reduces average training memory by up to 45.7\% without sacrificing performance. It also enables stable training of a 27B-parameter model for over 1,100 steps on a single eight-GPU node, where dense Adam runs out of memory, making previously memory-infeasible RL training practical. Code is available in the \href{https://github.com/skzhang1/labs-molt/tree/logra/examples/scripts/logra}{Molt library}.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Large Language Models
Memory Efficiency
Post-training
Innovation

Methods, ideas, or system contributions that make the work stand out.

Low-Rank Gradient Sketches
Reinforcement Learning
Memory Efficiency
Predicted-KL Step Control
Large Language Models