FlexLoop: Depth-Elastic Looped Policies for Adaptive Test-Time Computation in Deep RL

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inability of pretrained recurrent policies in deep reinforcement learning to adaptively adjust computation due to depth specialization. To overcome this limitation, we propose FlexLoop, a framework that employs a post-training paradigm based on adjacent-depth policy distillation to transform fixed-depth recurrent policies into depth-elastic ones. This approach endows shallower layers with reliable reasoning capabilities while preserving full-depth decision-making performance, thereby enabling dynamic, on-demand adjustment of recurrence depth during inference. Extensive experiments across 30 long-horizon tasks demonstrate that FlexLoop reduces the average recurrence depth by 43% and achieves a 1.34× wall-clock speedup, while maintaining performance comparable to that of full-depth models.
📝 Abstract
Looped architectures scale computation by reusing the same parameters across recurrent steps, and recent work shows that they substantially improve deep reinforcement learning policies on long-horizon tasks. Since recurrent depth directly controls computation, one may expect looped policies to naturally support elastic inference across recurrent depths. Surprisingly, we find that pretrained looped policies exhibit severe recurrent-depth specialization: reliable decisions are concentrated near the full trained depth, tying deployment computation to this depth even when less computation may suffice. Achieving depth elasticity, i.e., reliable decisions across recurrent depths with adaptive computation at deployment, therefore remains a key challenge. To address this, we propose FlexLoop, a novel post-training framework that converts pretrained fixed-depth looped policies into depth-elastic policies. FlexLoop keeps training on the original RL objective to preserve full-depth capability while performing adjacent-depth policy distillation to progressively transfer decision quality from deeper to shallower recurrent steps. The resulting policy supports reliable inference across recurrent depths and enables state-wise adaptive inference through recurrent-depth consistency. Experiments on $30$ online and offline long-horizon goal-conditioned environments show that FlexLoop preserves full-depth performance while making shallower depths effective. Keeping competitive performance, FlexLoop reduces average recurrent depth by up to $\bf{43\%}$ and achieves up to $\bf{1.34\times}$ wall-clock speedup in a stress test.
Problem

Research questions and friction points this paper is trying to address.

Deep Reinforcement Learning
Looped Policies
Depth Elasticity
Adaptive Test-Time Computation
Recurrent-Depth Specialization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Depth-Elastic Looped Policies
Test-Time Adaptive Computation
Adjacent-Depth Policy Distillation
Deep Reinforcement Learning
Recurrent-Depth Consistency
🔎 Similar Papers
X
Xun Wang
Institute for Interdisciplinary Information Sciences, Tsinghua University
R
Ruishuo Chen
Institute for Interdisciplinary Information Sciences, Tsinghua University
Yu Chen
Yu Chen
Tsinghua university
wearable robothuman-robot interaction
Z
Zhuoran Li
Institute for Interdisciplinary Information Sciences, Tsinghua University
Longbo Huang
Longbo Huang
Professor, IIIS, Tsinghua University, ACM Distinguished Scientist
Reinforcement Learning (RL)Deep RLMachine LearningStochastic NetworksPerformance Evaluation