Score
Designs, implements, and evaluates reinforcement-learning–based control policies for heating, ventilation, and air-conditioning systems—using methods such as PPO and deep RL—to map observed states to actions like supply-air setpoints and fan speeds. These controllers are built and analyzed to minimize energy consumption while enforcing thermal comfort constraints and providing robustness to stochastic disturbances.
Reinforcement learning (RL)-driven HVAC control faces challenges in cross-regional transferability across diverse urban climates, particularly concerning energy efficiency, indoor thermal comfort, and localized urban climate impacts. Method: This study establishes a multi-scale coupled simulation framework integrating RL, building energy modeling (EnergyPlus), urban canopy modeling (UCM), and meteorological data. It pioneers embedding HVAC RL policies within a climate feedback loop to quantify how ambient climate conditions govern reward function sensitivity and policy transferability, introducing the “inter-city learning” paradigm. Contribution/Results: Experiments reveal that hot cities achieve higher rewards under most trade-off configurations; cities with greater temperature variability exhibit stronger policy transferability; and RL deployment necessitates climate-specific evaluation. The work provides theoretical foundations and practical guidelines for climate-adaptive intelligent building control, advancing the design of resilient, energy-efficient urban infrastructure.
This study addresses the challenge of enabling building heating systems to efficiently respond to grid-side flexibility dispatch signals while maintaining indoor thermal comfort. The authors propose a reinforcement learning control framework that integrates the Deep Deterministic Policy Gradient (DDPG) algorithm with a building thermal dynamics model, augmented by a real-time adaptive safety filter to rigorously enforce compliance with system operator dispatch requirements. Simulation results demonstrate that the proposed method achieves up to 50% energy savings compared to rule-based controllers and outperforms pure reinforcement learning approaches. It reliably executes demand response instructions with 100% adherence while incurring only minor violations of thermal comfort constraints, thereby significantly enhancing both energy efficiency and dispatch reliability.
This study addresses the challenge of simultaneously optimizing energy efficiency and thermal comfort in building HVAC systems under nonlinear dynamics and stochastic loads, which conventional control strategies struggle to manage effectively. The authors propose an end-to-end deep reinforcement learning framework that directly controls the air handling unit (AHU) using the Proximal Policy Optimization (PPO) algorithm within a custom Python environment. The approach integrates a second-order RC thermal model with a CO₂ dynamic mass balance model to jointly optimize temperature and ventilation. A novel hierarchical airflow logic ensures indoor CO₂ concentrations remain at or below 1000 ppm, while an enthalpy-based economizer logic enables free cooling. Compared to traditional on–off control and a PID benchmark tuned via genetic algorithms, the proposed method demonstrates superior performance in both temperature stability and overall energy efficiency, validating its potential for intelligent building energy management.
This study addresses the persistent performance gap between simulation and real-world deployment of reinforcement learning (RL) in industrial energy systems, using a district heating network as a case study. The control task is formulated as a Markov decision process, and key challenges—including partial observability, action space design, reward function shaping, and sim-to-real transfer—are systematically analyzed. For the first time, this work provides a comprehensive empirical investigation of RL deployment barriers in an actual industrial setting, uncovering the root causes of performance discrepancies. Although the system achieved stable operation in practice, its measured performance fell significantly short of simulation results, thereby validating critical bottlenecks in real-world RL applications and establishing an empirical foundation for future co-optimization of algorithms and engineering design.
This study addresses critical challenges—scalability, safety, interpretability, and sample efficiency—in deploying reinforcement learning (RL) and model predictive control (MPC) for residential HVAC systems. For the first time, it conducts a one-month closed-loop comparative experiment in a real residential setting. A physics-informed MPC and a model-based RL (MBRL) approach, leveraging learned dynamic system models, are co-deployed with a heat pump to jointly optimize energy efficiency and thermal comfort. Results show RL achieves 22% energy savings—marginally exceeding MPC’s 20%—but MPC delivers superior energy efficiency at equivalent comfort levels. The study quantifies RL’s advantage in reducing modeling effort while empirically identifying, for the first time, key practical bottlenecks: unsafe policy initialization, instability during online adaptation, and actuation deviation. These findings establish a crucial benchmark and provide actionable design insights for operationalizing intelligent building control algorithms.
This study addresses the issue of excessive compressor cycling in residential heat pumps, which accelerates equipment wear—a factor commonly overlooked by existing reinforcement learning controllers that focus primarily on energy consumption and thermal comfort. To bridge this gap, the work explicitly incorporates compressor wear into the reward function and evaluates Soft Actor-Critic (SAC) and Proximal Policy Optimization (PPO) algorithms within the hydronic heat pump case of the BOPTEST platform. Results demonstrate that SAC autonomously learns a variable-speed continuous modulation strategy, achieving zero start-stop cycles while reducing thermal discomfort by 90.7% at the cost of only an 11.5% increase in operational expenditure, substantially outperforming conventional control approaches.
This study addresses the challenge of achieving efficient, safe, and adaptive control of water-cooled chillers and air-side systems in tropical commercial buildings. The authors propose the CQD-ERL controller, which introduces, for the first time in HVAC control, a context-aware quality-diversity (QD) evolutionary mechanism. By integrating evolutionary algorithms with the soft actor-critic (SAC) policy gradient method, the approach maintains a policy archive indexed jointly by operational context and behavioral descriptors, enabling the co-optimization of diverse, high-performing control policies. This framework supports dynamic selection of specialized policies tailored to prevailing weather conditions and load profiles. A deterministic safety shielding mechanism is incorporated to enforce critical constraints such as humidity levels and cooling tower approach temperature. Evaluated through full-year backtesting on a commercial building in Singapore, the proposed controller significantly outperforms the ASHRAE Guideline 36 baseline, achieving energy-efficient operation while rigorously maintaining safety constraints.
This work addresses the limited generalization of reinforcement learning (RL) policies in real-world deployment, where environmental dynamics shift, action and observation spaces vary, and control objectives change. To tackle this challenge, the authors introduce the first large-scale, physically realistic continuous control benchmark for HVAC control, built upon EnergyPlus. Leveraging a parameterized building generator, the benchmark systematically produces diverse building configurations and defines standardized tasks to evaluate key generalization capabilities—including objective adaptation, dynamic shifts, action space variations, and cross-domain transfer. The platform supports heterogeneous observation and action spaces and integrates with the Gymnasium interface alongside a unified evaluation protocol, thereby providing a robust foundation for advancing both building energy efficiency and the robustness of RL algorithms.