π€ AI Summary
This study addresses the joint capacity sizing and operational optimization of chillers and thermal energy storage (TES) in commercial HVAC systems by proposing a synergistic design framework that integrates deep reinforcement learning with life-cycle cost analysis. The system is modeled as a finite-horizon Markov decision process, and under a zero load-shedding constraint, a Deep Q-Network (DQN) is employed to optimize the chillerβs part-load ratio control policy within a constrained action space. The framework evaluates the 30-year life-cycle costs across various capacity combinations, effectively accounting for the pronounced asymmetry in capital expenditures between chillers and TES. The proposed approach achieves simultaneous optimization of both capacity allocation and operational strategy, yielding an optimal configuration of 700 units for the chiller and 1,500 units for the TES, which fully meets stochastic cooling demand while minimizing total life-cycle cost.
π Abstract
We study the joint operation and sizing of cooling infrastructure for commercial HVAC systems using reinforcement learning, with the objective of minimizing life-cycle cost over a 30-year horizon. The cooling system consists of a fixed-capacity electric chiller and a thermal energy storage (TES) unit, jointly operated to meet stochastic hourly cooling demands under time-varying electricity prices. The life-cycle cost accounts for both capital expenditure and discounted operating cost, including electricity consumption and maintenance. A key challenge arises from the strong asymmetry in capital costs: increasing chiller capacity by one unit is far more expensive than an equivalent increase in TES capacity. As a result, identifying the right combination of chiller and TES sizes, while ensuring zero loss-of-cooling-load under optimal operation, is a non-trivial co-design problem. To address this, we formulate the chiller operation problem for a fixed infrastructure configuration as a finite-horizon Markov Decision Process (MDP), in which the control action is the chiller part-load ratio (PLR). The MDP is solved using a Deep Q Network (DQN) with a constrained action space. The learned DQN RL policy minimizes electricity cost over historical traces of cooling demand and electricity prices. For each candidate chiller-TES sizing configuration, the trained policy is evaluated. We then restrict attention to configurations that fully satisfy the cooling demand and perform a life-cycle cost minimization over this feasible set to identify the cost-optimal infrastructure design. Using this approach, we determine the optimal chiller and thermal energy storage capacities to be 700 and 1500, respectively.