Hierarchical Reinforcement Learning with Stable Temporal Abstraction for Language Model Agents

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the instability of temporal abstraction in hierarchical language agents, which leads to premature subgoal replacement or stagnation. To this end, it proposes STAC, a method that innovatively formulates premature replanning and stale persistence as constraint costs for the first time. By employing Lagrangian optimization, these constraints exclusively modulate boundary policies without interfering with underlying reward mechanisms. Integrating hierarchical reinforcement learning with large language model agent techniques, STAC effectively enhances long-horizon task control. Evaluations on the ALFWorld and WebShop benchmarks demonstrate that STAC achieves up to a 23.5% improvement in success rate when instantiated with Qwen3 and Llama-3 backbones.
📝 Abstract
Hierarchical reinforcement learning improves long-horizon control by organizing primitive actions around persistent subgoals and assigning credit at multiple temporal scales. Recent hierarchical language agents bring these benefits to interactive tasks by explicitly separating subgoal planning from action execution. We observe, however, that an explicit hierarchy does not by itself determine how stable the resulting temporal abstraction is: the learned boundary policy may replace the subgoal almost every turn, making it effectively transient, or retain a subgoal after it has stopped being appropriate. We call this temporal abstraction instability. We propose Stable Temporal Abstraction via Constrained Optimization (STAC), a constrained boundary-policy optimization method that represents premature replanning and stale persistence as constraint costs. STAC applies the resulting Lagrangian costs only to the sampled boundary decision, leaving the underlying algorithm's rewards, critic targets, subgoal advantages, and primitive-action advantages unchanged. Across two backbones and two benchmarks, STAC improves success over a strong hierarchical baseline by $8.1$ and $7.9$ points on ALFWorld and WebShop with Qwen3-0.6B, and by $23.5$ and $15.8$ points with Llama-3.2-1B-Instruct.
Problem

Research questions and friction points this paper is trying to address.

Hierarchical Reinforcement Learning
Temporal Abstraction Instability
Language Model Agents
Boundary Policy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hierarchical Reinforcement Learning
Temporal Abstraction
Constrained Optimization
Language Model Agents
Boundary Policy