KV-streams for Efficient Compaction in Agentic Reinforcement Learning

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the GPU memory bottleneck caused by long contexts in agent reinforcement learning and the throughput limitations of conventional compression strategies. To this end, it proposes KV-streams, a plug-and-play method that achieves efficient context compression by streaming rather than resetting the KV cache. Furthermore, this work demonstrates for the first time that reinforcement learning alone can elicit the emergent ability to leverage streamed KV caches as recurrent states, without requiring additional supervision. Experimental results indicate that this strategy yields a 2.6× to 5× training speedup across three compression settings while preserving model performance.
📝 Abstract
Scaling the horizon of agentic LLMs is bottlenecked by the need to fit ever longer context traces in GPU memory. Context compaction has been the most popular mechanism to alleviate this issue, keeping GPU memory constant for a given trace. Unfortunately, most compaction strategies rely on prefilling the LLM context many times over, hindering training throughput. To alleviate this bottleneck and enable efficient trainable compaction, we propose KV-streams, a plug-and-play strategy compatible with any compaction strategy that substantially increases throughput while showing no evidence of hindering performance. KV-streams enable scalable compaction by streaming the KV cache forward rather than flushing it after each compaction. We show that KV-streams enable three different compaction strategies, achieving a 2.6 to 5x wall-clock speedup in training. Beyond efficiency, we find that the streamed KV cache can act as a recurrent state, carrying forward information that has long since disappeared from the context. Specifically, in a controlled setting we show that, contrary to prior work, RL alone is all that is needed for this behavior to emerge. Overall, we show KV-streams to be an efficient and lightweight plug-and-play addition to any post-training pipeline.
Problem

Research questions and friction points this paper is trying to address.

Agentic Reinforcement Learning
Context Compaction
KV Cache
Training Throughput
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

KV-streams
Context Compaction
Agentic Reinforcement Learning
KV Cache Streaming
Recurrent State
🔎 Similar Papers
No similar papers found.
E
Emiliano Penaloza
Mila
D
Dane Malenfant
McGill University
Dheeraj Vattikonda
Dheeraj Vattikonda
MILA - Quebec Artificial Intelligence Institute
Reinforcement Learning
Roger Creus Castanyer
Roger Creus Castanyer
Mila/University of Montreal
Reinforcement LearningFoundation Models
Siddarth Venkatraman
Siddarth Venkatraman
Mila, University of Montreal
Artificial IntelligenceRobotics
Abhay Puri
Abhay Puri
Applied Research Scientist, ServiceNow Research
Agent SecurityLarge Language ModelsComputer VisionMultiModal Foundational Models
Jonathan Light
Jonathan Light
RPI PhD
Decision making under uncertaintyfoundation modelsreinforcement learning
M
Matthew James Sargent
University College London, University of London
A
Augustine N. Mavor-Parker
Vmax
M
Massimo Caccia
Cohere
Lucas Caccia
Lucas Caccia
Microsoft Research
Deep LearningContinual LearningNatural Language Processing
Glen Berseth
Glen Berseth
Assitant Professor - Université de Montréal
Reinforcement LearningRoboticsDeep LearningMachine Learning
E
Esmeralda S. Whitammer
Edinburgh University
Alessandro Sordoni
Alessandro Sordoni
Microsoft Research
Artificial IntelligenceInformation RetrievalDeep Learning
Minseon Kim
Minseon Kim
Microsoft Research
AI SafetyRobustnessRepresentation learning
Marc-Alexandre Côté
Marc-Alexandre Côté
Microsoft Research
Deep LearningNatural Language UnderstandingReinforcement LearningText-based Games
Laurent Charlin
Laurent Charlin
Associate Professor, HEC Montréal & Mila, Canada CIFAR AI Chair
Machine LearningArtificial Intelligence
Guillaume Lajoie
Guillaume Lajoie
Professor, Mila & Université de Montréal
AIdynamical systemscomputation neurosciencenetwork dynamicsmachine learning theory