BAPR: Bayesian amnesic piecewise-robust reinforcement learning for non-stationary continuous control

📅 2026-05-15
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of balancing performance during stable phases and safety during abrupt shifts in piecewise-stationary non-stationary continuous control environments, where conventional robust reinforcement learning methods often fail. The authors propose an adaptive framework that integrates Bayesian Online Change Point Detection (BOCD) with a robust ensemble variant of Soft Actor-Critic (SAC). Central to this approach is a belief-weighted BAPR operator endowed with γ-contraction properties, which rapidly increases policy conservatism following a detected change and gradually relaxes it as confidence in stationarity is restored. The paper formally characterizes, for the first time, the critical conditions under which belief-dependent Q-functions violate contraction and provides component-wise verifiable error budgets. A context-conditioning module obviates the need for mode labels at deployment. All theoretical claims are machine-verified in Lean4 (1,145 lines, 22 theorems), achieving O(log(1/δ)) detection delay and ensuring effective post-shift policy recovery.
📝 Abstract
Real-world control systems frequently operate under \emph{piecewise stationary} conditions, where dynamics remain stable for extended periods before undergoing abrupt regime changes. Standard robust RL methods face a fundamental dilemma: a globally conservative policy wastes performance during stable periods, while a locally adaptive policy risks catastrophic failure when the regime changes undetected. We propose \textbf{BAPR} (Bayesian Amnesic Piecewise-Robust SAC), which unifies Bayesian Online Change Detection (BOCD) with robust ensemble RL. The BAPR operator -- a convex combination of mode-conditional Bellman operators weighted by a frozen belief distribution -- is a $\gamma$-contraction. A complementary counterexample, machine-verified in Lean~4, establishes a \emph{sharp boundary}: when beliefs depend on the Q-function, the contraction factor becomes $\gamma + \lambda\Delta$ (where $\Delta$ is the mode reward gap), and contraction fails exactly when $\gamma + \lambda\Delta \geq 1$. We derive a \emph{component-wise} formal error budget for the abstract operator -- every component machine-verified -- bounding post-switch recovery; the budget applies to the abstract mode-mixture operator and inherits to the implemented shared-critic algorithm only through the frozen-parameter design intuition. All results are formally verified with no \texttt{sorry} (1,145 lines across 3 Lean~4 files, 22 machine-verified theorems). BOCD drives an adaptive conservatism mechanism: the policy becomes maximally conservative after detected change-points and smoothly relaxes as confidence grows, with detection delay $O(\log(1/\delta))$. A context-conditioning module trained via RMDM loss provides mode-aware representations from simulator-provided mode IDs at training time and requires no mode labels at deployment.
Problem

Research questions and friction points this paper is trying to address.

non-stationary
piecewise stationary
robust reinforcement learning
regime change
continuous control
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bayesian Online Change Detection
Robust Reinforcement Learning
Contraction Mapping
Formal Verification
Non-stationary Control
💼 Related Jobs
No related jobs found.