🤖 AI Summary
This study investigates the causal relationship between information acquisition and commitment behavior in safe adaptive control: specifically, whether a controller can distinguish between system models requiring distinct policies through informative experiments within finite time under unified safety constraints. To address this, the work introduces the notion of “pre-commitment information,” quantifies observable information via Kullback–Leibler divergence, and integrates causal analysis, semidefinite programming, and linear Gaussian system theory to establish fundamental limits on information gathering under safety constraints. The main contributions include proving that in linear quadratic regulation, if the oracle gap is Ω(T), then every uniformly safe policy incurs linear regret; furthermore, the paper provides sufficient conditions for recoverability together with corresponding upper-bound certificates.
📝 Abstract
Safe adaptive control is online adaptation under a safety guarantee on the learning trajectory itself. The controller may use any causal, history-dependent rule and act differently across environments as data arrive. Only its safety guarantee is uniform: the same rule must satisfy it under every initially plausible model. Performance is measured against a safe oracle that knows the realized model. Many finite-time analyses assume persistent excitation of the uniformly safe closed loop, so the data distinguish every pair of models requiring different control decisions. Under that assumption, feasibility is already settled; only the rate remains. We ask instead: Do the safety constraints permit such an informative experiment at all?
While an alternative remains plausible, the controller must preserve a safe continuation under it. We call the first action that forecloses such a continuation commitment. Chance safety allows commitment only on an event rare under the alternative, and the evidence must arrive beforehand: the observation generated by the committing action is too late. We define precommitment information as the KL divergence between learner-visible laws stopped before commitment.
Our main result is a causal reduction. The commitment rule determines (1) the probability that safety permits commitment under the alternative, (2) the target-side cost of remaining noncommittal, (3) and the information available when the decision is made. Bounded precommitment information therefore leaves a fixed fraction of the oracle gap unavoidable. If the gap is Ω(T), every uniformly safe policy has linear regret. We establish the obstruction in a constrained linear system with quadratic regulation cost. We also prove recovery in special cases and derive semidefinite upper certificates for deterministic linear-Gaussian systems.