Learning Multiple Timescales for Goal-Conditioned Reinforcement Learning

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of discount-induced value signal decay and the inability of fixed temporal abstraction to simultaneously capture local and global dynamics in offline goal-conditioned reinforcement learning for long-horizon tasks. We propose Generalized Implicit Temporal Abstraction (GITA), which conditions a single value function on multi-scale k-values and aggregates them via advantage-weighted regression. To our knowledge, this is the first approach to explicitly model the temporal abstraction trade-off, preserving both local precision and long-range discriminability without requiring the selection of a single k-value. Evaluated on the OGBench benchmark, GITA improves average success rates by 25 percentage points over HIQL and by 7 percentage points over OTA, the strongest fixed-k baseline.
📝 Abstract
Existing approaches to offline goal-conditioned reinforcement learning (GCRL) struggle with long-horizon tasks. Discounting shrinks value differences between distant states until they fall below the function approximation error, leaving the agent with no signal for ranking states. Temporal abstraction, which treats k environment steps as a single transition, restores this signal at long range, but no single fixed k suits all state-goal distances: large k preserves value differences across long temporal distances while collapsing distinctions between nearby states, and small k does the reverse. We make this trade-off explicit and introduce Generalized Implicit Temporal Abstraction (GITA), which conditions a single value function on k. GITA trains one policy by aggregating advantage-weighted supervision across multiple k values, so scales assigning larger positive advantages to a state-goal pair contribute more strongly to its update. GITA does not need to choose between local resolution and long-range signal; it retains both without committing to a single k. On OGBench, GITA outperforms a broad range of offline GCRL baselines, raising average success rate across all tasks by 25 percentage points (73% relative improvement) over HIQL. It also improves over the strongest fixed-k method, OTA, by 7 percentage points (14% relative).
Problem

Research questions and friction points this paper is trying to address.

Goal-Conditioned Reinforcement Learning
Long-Horizon Tasks
Temporal Abstraction
Offline RL
Innovation

Methods, ideas, or system contributions that make the work stand out.

Goal-Conditioned Reinforcement Learning
Temporal Abstraction
Offline Reinforcement Learning
Generalized Implicit Temporal Abstraction (GITA)
Advantage-Weighted Supervision
🔎 Similar Papers
No similar papers found.
Pedro Robles Dutenhefner
Pedro Robles Dutenhefner
Pesquisador Universidade Federal de Minas Gerais
D
Dikshant Shehmar
University of Alberta, Alberta Machine Intelligence Institute (Amii)
W
Wagner Meira Jr.
Universidade Federal de Minas Gerais
M
Marlos C. Machado
University of Alberta, Alberta Machine Intelligence Institute (Amii), Canada CIFAR AI Chair