Efficient Bayes-Adaptive Reinforcement Learning with Temporal Logic Specifications

📅 2026-09-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种基于贝叶斯自适应马尔可夫决策过程和限界确定性布希自动机同步的强化学习算法,用于在未知环境中高效合成满足线性时序逻辑规范的策略。
📝 Abstract
We present a novel end-to-end model-based Reinforcement Learning (RL) algorithm for efficient policy synthesis under given Linear Temporal Logic (LTL) specifications (e.g., safety or reachability) in unknown environments. To do so, a Limit-Deterministic B{ü}chi Automaton (LDBA) representation of the LTL task is synchronised with a Bayes-Adaptive Markov Decision Process (BAMDP) representation of the environment, which allows us to leverage an enhanced exploration-exploitation trade-off that is achieved via Bayesian RL, as opposed to traditional non-Bayesian approaches. We further propose a novel Bayes-Adaptive Monte-Carlo Planning (BAMCP) algorithm to allow for approximate Bayes-optimal strategy synthesis in the synchronised BAMDP construct. A range of finite- and infinite-horizon task experiments demonstrate the effectiveness of our approach in terms of both property satisfaction and sample efficiency, when compared to traditional model-free approaches. Additional ablation studies also successfully highlight the value of the novel BAMCP algorithm in comparison to classical BAMCP for LTL task satisfaction. Finally, we also showcase a successful application of our approach for \textit{cautious} RL, namely to reduce the number of task violations incurred during policy training.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Linear Temporal Logic
Bayesian RL
BAMDP
Policy Synthesis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bayes-Adaptive Reinforcement Learning
Linear Temporal Logic
Bayes-Adaptive Monte-Carlo Planning
Limit-Deterministic B{\
💼 Related Jobs
No related jobs found.