Q-Learning for Reachability in MEC-Free MDPs

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the prohibitive memory overhead of O(|S|²|A|) in existing reinforcement learning approaches for reachability specifications, which inherently rely on transition probability estimation. To overcome this limitation, this work proposes Quasar, a novel algorithm grounded in classical Q-learning and temporal difference updates. Operating on MEC-free MDPs, Quasar establishes the first model-free reachability learning framework with asymptotic convergence guarantees. The primary contribution lies in achieving convergence to optimal policies without estimating transition probabilities, thereby reducing memory complexity to O(|S||A|). Furthermore, empirical evaluations on standard benchmarks demonstrate that Quasar attains convergence using orders of magnitude fewer samples. These advances significantly facilitate the practical deployment of specification-guided reinforcement learning.
📝 Abstract
Reinforcement learning (RL) for reachability specifications is fundamental to sequential decision-making. Prior work establishes asymptotic convergence to optimal policies, but only through model-based methods that must explicitly estimate the transition probabilities of the underlying Markov Decision Process (MDP). We present Quasar, the first model-free algorithm with asymptotic guarantees for reachability on the fragment of MDPs free of non-terminal maximal end components (MECs), a building block to which every MDP reduces by the standard MEC quotient. Our algorithm follows the classical Q-learning approach, using temporal-difference updates to converge to an optimal policy without ever learning the transition probabilities. The resulting learner reduces the memory footprint from the O(|S|^2|A|) that model-based methods require to O(|S||A|). On the standardized Quantitative Verification Benchmark Set, our algorithm converges to the optimal policy with orders of magnitude fewer samples than the previous model-based state-of-the-art. Together these results are a concrete step toward the practical deployment of reachability learning and, with it, of specification-guided RL.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Reachability
Model-Free
Markov Decision Process
Q-Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Model-free reinforcement learning
Q-learning
Reachability
Markov Decision Process
Sample efficiency
🔎 Similar Papers