🤖 AI Summary
This study addresses the challenge of ensuring safety when embedding industrial control models into deployed systems, particularly regarding safety constraints in pressurized water reactor (PWR) load-following operations. To this end, a physics-decomposition-based structured representation learning method is proposed. This approach constructs a multi-timescale separation embedding architecture that maps variables across different timescales into independent latent spaces to emulate expert policies. It further enables hybrid deployment by integrating behavioral cloning with nonlinear model predictive control (NMPC). Experimental results demonstrate that the proposed method significantly improves the accuracy and feasibility of long-horizon trajectories, achieving fully feasible solutions with near-optimal costs while reducing computation time by approximately 15% compared to the expert controller.
📝 Abstract
Learned models for industrial control are usually judged by aggregate accuracy, but accuracy at the component level does not guarantee safety once it is embedded in the system it is meant to serve. We study this gap on a behavior-cloning task: imitating an expert Nonlinear Model Predictive Control (NMPC) policy for load-following of a Pressurized Water Reactor (PWR), an industrial system with tight safety constraints. We propose a structured architecture encoding variables from each timescale into separate latent spaces, reflecting the physical decomposition of the system, before training a controller to imitate the expert on the product latent space. On long-horizon rollouts, separated embeddings improve both accuracy and feasibility compared with a shared-embedding baseline. Sensitivity analysis further shows that our model yields interpretable representations aligned with the system's physics. However, standalone deployment still leaves several percent of trajectories infeasible regardless of the architecture. Using our method to warmstart the NMPC optimizer rather than acting standalone, we recover full feasibility and near-optimal cost while still cutting computation time by $\sim$15% relative to the expert controller, and even more for abrupt operating changes.