🤖 AI Summary
This work investigates how intelligent agents can dynamically balance fast reactive control against slower yet more robust deliberative planning to achieve both efficiency and performance. To this end, the authors propose a learnable meta-reasoning controller trained via reinforcement learning that adaptively triggers planning based on an uncertainty score derived from the reactive policy. The framework integrates reinforcement learning, imitation learning, model-based planning, and uncertainty estimation, enabling the agent to progressively shift toward purely reactive control as the reactive policy improves. Experiments in motion planning and navigation tasks demonstrate that the agent accurately discerns when to rely on reactive responses versus when to invoke planning, dynamically optimizing computational resource allocation throughout training and yielding a flexible, efficient decision-making architecture.
📝 Abstract
It has long been recognized that humans have the ability to switch between fast, reactive decision-making and slower, deliberative planning. In this paper, we study the question of how to learn this ability, known as meta-reasoning, in artificial agents. We model reactive decision-making as a policy that directly maps state observations to actions. Such policies can be trained with reinforcement learning (RL) or imitation learning, but may generalize poorly outside of their training distribution. Alternatively, model-based decision-time planning is more likely to produce good actions across a broader set of states but requires additional computation time, which delays acting. In this work, we introduce an RL method for training a meta-reasoning policy that allocates computation by conditioning on a reactive-policy uncertainty score. This score enables it to predict when the reactive policy is likely to perform poorly and when planning is needed. We conduct an empirical study on motion planning and navigation environments, showing that this design enables the meta-reasoning policy to learn when the reactive policy provides a good-enough action versus when decision-time planning is needed. Additionally, we show that our design enables the meta-agent to shift toward fully reactive control as the reactive policy improves.