Score
Designs, implements, or evaluates a meta-controller — a high-level policy that selects among temporally-extended actions (options) and controls their initiation, termination, and sequencing within the options framework. This includes learning the option-selection policy and associated initiation/termination mechanisms as part of options-based hierarchical reinforcement learning.
This work investigates how intelligent agents can dynamically balance fast reactive control against slower yet more robust deliberative planning to achieve both efficiency and performance. To this end, the authors propose a learnable meta-reasoning controller trained via reinforcement learning that adaptively triggers planning based on an uncertainty score derived from the reactive policy. The framework integrates reinforcement learning, imitation learning, model-based planning, and uncertainty estimation, enabling the agent to progressively shift toward purely reactive control as the reactive policy improves. Experiments in motion planning and navigation tasks demonstrate that the agent accurately discerns when to rely on reactive responses versus when to invoke planning, dynamically optimizing computational resource allocation throughout training and yielding a flexible, efficient decision-making architecture.
This work addresses the challenge of ensuring both safety and sample-efficient adaptation during test-time task learning in meta-reinforcement learning. It proposes the first constrained meta-RL algorithm that simultaneously guarantees provable safety at test time and achieves near-optimal sample complexity. The approach learns a general-purpose prior during meta-training and refines the policy under explicit safety constraints during test-time adaptation. Theoretical analysis demonstrates that the method converges to a near-optimal policy with a sample complexity that matches the established lower bound, while rigorously satisfying safety requirements throughout the learning process.
研究通过层次潜因器探讨了高级控制器何时应保留或修改低级过程策略的问题,发现状态依赖并不等同于决策价值。
Deep reinforcement learning (DRL) suffers from low data efficiency and poor generalization across tasks. Meta-reinforcement learning (Meta-RL) addresses this by treating algorithm design itself as a learning problem, enabling agents to rapidly adapt to new tasks with few samples drawn from a task distribution. This paper introduces the first unified classification framework for Meta-RL, characterized along two orthogonal dimensions: (i) whether the task distribution is explicitly modeled, and (ii) whether the per-task learning budget is constrained. We systematically survey problem formulations, core paradigms, and representative algorithms—including MAML-based, RNN-based, contextual, and Bayesian approaches. Our contributions are threefold: (i) clarifying the field’s evolutionary trajectory and identifying key open challenges; (ii) constructing the first structured pedagogical guide; and (iii) advancing Meta-RL toward becoming a standard tool in the DRL toolkit—now widely adopted for both education and practical onboarding.
To address low sample efficiency in meta-testing due to the absence of reward signals, this paper proposes UMCNP, an unsupervised meta-reinforcement learning framework. UMCNP integrates policy gradient optimization with task inference, enabling implicit environment dynamics modeling and adaptive policy optimization from a single trajectory of an unseen task. Its key contributions are threefold: (1) decoupling policy learning from task inference to support offline reuse of meta-training data; (2) employing Conditional Neural Processes (CNPs) for unsupervised task representation learning; and (3) combining parameterized policy gradients with a model-predictive control–inspired self-generated rollout mechanism. Evaluated on benchmarks—including 2D point navigation, biased-sensor CartPole, and dynamics-randomized Walker—UMCNP reduces meta-test sample requirements by over 50% while significantly improving few-shot adaptation performance.
This work addresses the instability and lack of interpretability in directly using large language models (LLMs) for trading decisions, as well as the inadequacy of single fixed strategies in dynamic markets. The authors propose an executable decision-making paradigm that adaptively selects optimal modules from a programmable strategy library and generates actions conditioned on market states, rather than emitting raw trading signals. Their approach leverages simulation-based supervision, converting simulated performance into state–strategy paired data to enable efficient and interpretable strategy selection. Experiments across multi-asset and commodity trading environments demonstrate that fine-tuned models ranging from 0.8B to 9B parameters significantly outperform fixed strategies, end-to-end LLMs, and API-based agents—with even smaller models surpassing more powerful API-driven counterparts.
This study addresses the challenge of rapidly adapting policies to novel environments and interactive tasks in multi-agent systems by proposing a bilevel optimization-based meta multi-agent reinforcement learning framework. The core innovation lies in defining the concept of "meta-Nash equilibrium" and rigorously establishing sufficient conditions for its equivalence to the stationary points of gradient-based game dynamics algorithms. Methodologically, this work integrates Markov game modeling with meta-reinforcement learning mechanisms to enable rapid policy adaptation across a distribution of games. Experimental evaluations on autonomous driving tasks demonstrate that the proposed approach achieves superior adaptation efficiency compared to pretrained baselines, thereby validating the effectiveness of the theoretical framework.
This work addresses the limitations of existing meta-reinforcement learning approaches, which often couple task inference with policy execution, leading to ambiguous task semantics, low sample efficiency, and poor knowledge transfer across heterogeneous agents. To overcome these issues, the authors propose a decoupled meta-knowledge reuse framework that learns task-level knowledge on simplified dynamical surrogates. Task structures are organized via a Bayesian nonparametric prior, and knowledge is transferred to diverse agents through a semantic-magnitude interface combined with a lightweight temporal adapter. This design enables frozen-task knowledge reuse, substantially improving cross-embodiment transfer efficiency. Experiments demonstrate that the method achieves performance comparable to state-of-the-art baselines using only approximately 23.8% of the interaction data across multiple locomotion agents, while reducing final tracking errors by 94.75%–99.79%.
Existing hierarchical decision-making approaches often struggle to simultaneously satisfy constraints and maintain computational efficiency due to misalignment between low-level policies and high-level objectives. This work proposes a principled inverse optimization–based hierarchical framework that, for the first time, systematically constructs structured low-level optimization problems from expert demonstrations, thereby aligning high-level task abstractions with low-level decision-making. By integrating inverse optimization, hierarchical reinforcement learning, and optimal control, the method achieves both interpretability and computational efficiency. Empirical evaluations on resource allocation and obstacle avoidance tasks demonstrate that the approach significantly outperforms end-to-end reinforcement learning, learning-augmented optimal control, and existing hierarchical methods, achieving state-of-the-art performance in both decision quality and computational speed.
This study addresses the difficulty of pretrained robot policies in adapting to unknown physical environments due to their reliance on sparse rewards and limited use of geometric dynamic feedback. To overcome this, we propose SCOUT, a framework that couples action prediction with forward dynamics models to construct a shared belief latent space. By leveraging dynamics-aware meta-learning, it establishes a bilevel architecture comprising an inner loop for belief updating and an outer loop for action optimization. The core innovation lies in utilizing action-outcome feedback to update internal dynamics beliefs online, enabling rapid and robust adaptation without sparse rewards while avoiding catastrophic forgetting. Experiments demonstrate that SCOUT significantly accelerates online adaptation in simulated manipulation benchmarks and successfully validates robust sim-to-real transfer capabilities.