🤖 AI Summary
This work addresses the challenge of balancing performance during stable phases and safety during abrupt shifts in piecewise-stationary non-stationary continuous control environments, where conventional robust reinforcement learning methods often fail. The authors propose an adaptive framework that integrates Bayesian Online Change Point Detection (BOCD) with a robust ensemble variant of Soft Actor-Critic (SAC). Central to this approach is a belief-weighted BAPR operator endowed with γ-contraction properties, which rapidly increases policy conservatism following a detected change and gradually relaxes it as confidence in stationarity is restored. The paper formally characterizes, for the first time, the critical conditions under which belief-dependent Q-functions violate contraction and provides component-wise verifiable error budgets. A context-conditioning module obviates the need for mode labels at deployment. All theoretical claims are machine-verified in Lean4 (1,145 lines, 22 theorems), achieving O(log(1/δ)) detection delay and ensuring effective post-shift policy recovery.
📝 Abstract
Real-world control systems frequently operate under \emph{piecewise stationary} conditions, where dynamics remain stable for extended periods before undergoing abrupt regime changes. Standard robust RL methods face a fundamental dilemma: a globally conservative policy wastes performance during stable periods, while a locally adaptive policy risks catastrophic failure when the regime changes undetected. We propose \textbf{BAPR} (Bayesian Amnesic Piecewise-Robust SAC), which unifies Bayesian Online Change Detection (BOCD) with robust ensemble RL. The BAPR operator -- a convex combination of mode-conditional Bellman operators weighted by a frozen belief distribution -- is a $\gamma$-contraction. A complementary counterexample, machine-verified in Lean~4, establishes a \emph{sharp boundary}: when beliefs depend on the Q-function, the contraction factor becomes $\gamma + \lambda\Delta$ (where $\Delta$ is the mode reward gap), and contraction fails exactly when $\gamma + \lambda\Delta \geq 1$. We derive a \emph{component-wise} formal error budget for the abstract operator -- every component machine-verified -- bounding post-switch recovery; the budget applies to the abstract mode-mixture operator and inherits to the implemented shared-critic algorithm only through the frozen-parameter design intuition. All results are formally verified with no \texttt{sorry} (1,145 lines across 3 Lean~4 files, 22 machine-verified theorems). BOCD drives an adaptive conservatism mechanism: the policy becomes maximally conservative after detected change-points and smoothly relaxes as confidence grows, with detection delay $O(\log(1/\delta))$. A context-conditioning module trained via RMDM loss provides mode-aware representations from simulator-provided mode IDs at training time and requires no mode labels at deployment.