🤖 AI Summary
This study addresses the challenge that the distinguishability assumptions underlying offline guarantees in non-Markovian decision processes are difficult to verify and prone to failure. Drawing upon finite automata theory and Bayesian inference, we prove the posterior odds invariance of observationally equivalent models, thereby providing the first precise characterization of indistinguishability. Based on this theoretical foundation, we propose PEC, a linear-time decision algorithm, and formally verify its core theorems using Lean 4. Experimental results demonstrate that conventional assumptions frequently fail in practice, whereas the PEC algorithm effectively restores distinguishability, ensuring that the theoretical guarantees for offline policy evaluation remain valid.
📝 Abstract
Non-Markovian environments are often modeled as Regular Decision Processes (RDPs), where dynamics depend on the interaction history through a finite automaton. Existing offline guarantees for RDPs rely on a distinguishability assumption on the behaviour policy but provide no means of verifying it. When the assumption is violated, distinct models may explain the data equally well. We study when data collected under a fixed behaviour policy can distinguish two candidate RDPs. We prove that the posterior odds between observationally equivalent candidates remain equal to the prior odds at every sample size, even when the policy visits every automaton state, and verify both results formally in Lean 4. We then characterize this equivalence exactly and derive PEC, an algorithm that decides it in time linear in the size of the product automaton. The distinguishability assumption of prior work fails on three of our four test environments, and the experiment identified by PEC restores it in each case.