Computationally efficient safe exploration in reinforcement learning

📅 2026-09-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种基于Nadaraya-Watson估计器的轻量级算法,以解决强化学习中安全探索的问题,特别是在约束马尔可夫决策过程中。
📝 Abstract
Reinforcement learning in real-life applications requires safety guarantees during exploration. Typical reinforcement learning algorithms do not provide such guarantees, and many modifications that do rely on Gaussian processes (GPs), which have a large computational cost. We propose a computationally lightweight algorithm based on the Nadaraya-Watson estimator that safely explores and optimizes constrained Markov decision processes (MDPs). Our algorithm, \textsc{CoLSafe-MDP}, uses an estimator that scales in constant-time with bounds on the estimates, a significant improvement from its GP-based counterparts that scale cubically with the number of data points. We then evaluate its performance in a grid-based environment and on observational Martian terrain data.
Problem

Research questions and friction points this paper is trying to address.

reinforcement learning
safety guarantees
exploration
Gaussian processes
computational cost
Innovation

Methods, ideas, or system contributions that make the work stand out.

Nadaraya-Watson estimator
computationally efficient
safe exploration
constrained MDPs