Improving planning and MBRL with temporally-extended actions

📅 2025-05-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Discrete-time modeling of continuous-time systems in model-based reinforcement learning (MBRL) and trajectory optimization leads to excessive planning steps, high computational overhead, and cumulative model prediction errors. Method: We propose the Temporal-Expansion Action (TEA) framework, which treats action duration as an optimizable variable to explicitly control the decision-making timescale. TEA is the first approach to jointly optimize action duration within MBRL and trajectory optimization; it employs a multi-armed bandit to adaptively select duration ranges and decouples deep primitive-action horizons from shallow planning depths. Contribution/Results: Experiments demonstrate that TEA significantly accelerates planning, improves solution quality, resolves convergence failures of standard methods on multiple benchmark tasks, reduces model training time, and mitigates error accumulation—thereby enhancing both efficiency and robustness of continuous-control planning.

Technology Category

Planning, Routing, and Scheduling: Temporal PlanningMultiagent Systems: Multiagent PlanningHumans and AI: Human-Aware Planning and Behavior Prediction

Application Category

User Modeling, Personalization and Recommendation: Studies of user behavior, including longitudinal effects of personalized systemsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingResponsible Web: Machine-in-the-loop, human agency and autonomy
📝 Abstract
Continuous time systems are often modeled using discrete time dynamics but this requires a small simulation step to maintain accuracy. In turn, this requires a large planning horizon which leads to computationally demanding planning problems and reduced performance. Previous work in model free reinforcement learning has partially addressed this issue using action repeats where a policy is learned to determine a discrete action duration. Instead we propose to control the continuous decision timescale directly by using temporally-extended actions and letting the planner treat the duration of the action as an additional optimization variable along with the standard action variables. This additional structure has multiple advantages. It speeds up simulation time of trajectories and, importantly, it allows for deep horizon search in terms of primitive actions while using a shallow search depth in the planner. In addition, in the model based reinforcement learning (MBRL) setting, it reduces compounding errors from model learning and improves training time for models. We show that this idea is effective and that the range for action durations can be automatically selected using a multi-armed bandit formulation and integrated into the MBRL framework. An extensive experimental evaluation both in planning and in MBRL, shows that our approach yields faster planning, better solutions, and that it enables solutions to problems that are not solved in the standard formulation.
Problem

Research questions and friction points this paper is trying to address.

Addressing computational demands in discrete-time planning with continuous systems
Reducing planning horizon by optimizing action durations directly
Improving MBRL performance by minimizing compounding model errors
Innovation

Methods, ideas, or system contributions that make the work stand out.

Controls continuous decision timescale with temporally-extended actions
Treats action duration as additional optimization variable
Uses multi-armed bandit for automatic action duration selection
💼 Related Jobs
No related jobs found.