🤖 AI Summary
To address coordinated control of multiple subprocesses in cyber-physical systems, this paper proposes a dual-timescale hierarchical decentralized control architecture: a global controller—formulated as an infinite-horizon discounted Markov decision process (MDP)—optimizes overall performance under budget constraints; meanwhile, (N) local controllers, each modeled as an MDP, autonomously make decisions under either constrained optimization (COpt) or unconstrained optimization (FOpt) frameworks. We establish, for the first time, a rigorous theoretical connection between COpt and FOpt: proving existence of optimal policies, deriving bounds on the difference between their optimal value functions, and characterizing equivalence conditions. We further prove that static deterministic optimal policies exist and identify precise budget–cost matching conditions under which COpt and FOpt become equivalent. These results provide a sound theoretical foundation and principled design guidelines for decentralized, scalable, federated autonomous control.
📝 Abstract
This paper presents a two-timescale hierarchical decentralized architecture for control of Cyber-Physical Systems. The architecture consists of $N$ independent sub-processes, a global controller, and $N$ local controllers, each formulated as a Markov Decision Process (MDP). The global controller, operating at a slower timescale optimizes the infinite-horizon discounted cumulative reward under budget constraints. For the local controllers, operating at a faster timescale, we propose two different optimization frameworks, namely the COpt and FOpt. In the COpt framework, the local controller also optimizes an infinite-horizon MDP, while in the FOpt framework, the local controller optimizes a finite-horizon MDP. The FOpt framework mimics a federal structure, where the local controllers have more autonomy in their decision making. First, the existence of stationary deterministic optimal policies for both these frameworks is established. Then, various relationships between the two frameworks are studied, including a bound on the difference between the two optimal value functions. Additionally, sufficiency conditions are provided such that the two frameworks lead to the same optimal values.