rl hvac control

Designs, implements, and evaluates reinforcement-learning–based control policies for heating, ventilation, and air-conditioning systems—using methods such as PPO and deep RL—to map observed states to actions like supply-air setpoints and fan speeds. These controllers are built and analyzed to minimize energy consumption while enforcing thermal comfort constraints and providing robustness to stochastic disturbances.

rlhvaccontrol

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.16
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Reinforcement Learning (RL) Meets Urban Climate Modeling: Investigating the Efficacy and Impacts of RL-Based HVAC Control

May 11, 2025
JY
Junjie Yu
🏛️ The University of Manchester | NSF National Center for Atmospheric Research | China University of Petroleum

Reinforcement learning (RL)-driven HVAC control faces challenges in cross-regional transferability across diverse urban climates, particularly concerning energy efficiency, indoor thermal comfort, and localized urban climate impacts. Method: This study establishes a multi-scale coupled simulation framework integrating RL, building energy modeling (EnergyPlus), urban canopy modeling (UCM), and meteorological data. It pioneers embedding HVAC RL policies within a climate feedback loop to quantify how ambient climate conditions govern reward function sensitivity and policy transferability, introducing the “inter-city learning” paradigm. Contribution/Results: Experiments reveal that hot cities achieve higher rewards under most trade-off configurations; cities with greater temperature variability exhibit stronger policy transferability; and RL deployment necessitates climate-specific evaluation. The work provides theoretical foundations and practical guidelines for climate-adaptive intelligent building control, advancing the design of resilient, energy-efficient urban infrastructure.

Assessing impacts of RL strategies on indoor and urban climatesEvaluating RL-based HVAC control efficacy in diverse climatesExploring transferability of RL strategies across different cities

This study addresses the challenge of enabling building heating systems to efficiently respond to grid-side flexibility dispatch signals while maintaining indoor thermal comfort. The authors propose a reinforcement learning control framework that integrates the Deep Deterministic Policy Gradient (DDPG) algorithm with a building thermal dynamics model, augmented by a real-time adaptive safety filter to rigorously enforce compliance with system operator dispatch requirements. Simulation results demonstrate that the proposed method achieves up to 50% energy savings compared to rule-based controllers and outperforms pure reinforcement learning approaches. It reliably executes demand response instructions with 100% adherence while incurring only minor violations of thermal comfort constraints, thereby significantly enhancing both energy efficiency and dispatch reliability.

building heating controldemand-side flexibilityenergy efficiency

This study addresses the challenge of simultaneously optimizing energy efficiency and thermal comfort in building HVAC systems under nonlinear dynamics and stochastic loads, which conventional control strategies struggle to manage effectively. The authors propose an end-to-end deep reinforcement learning framework that directly controls the air handling unit (AHU) using the Proximal Policy Optimization (PPO) algorithm within a custom Python environment. The approach integrates a second-order RC thermal model with a CO₂ dynamic mass balance model to jointly optimize temperature and ventilation. A novel hierarchical airflow logic ensures indoor CO₂ concentrations remain at or below 1000 ppm, while an enthalpy-based economizer logic enables free cooling. Compared to traditional on–off control and a PID benchmark tuned via genetic algorithms, the proposed method demonstrates superior performance in both temperature stability and overall energy efficiency, validating its potential for intelligent building energy management.

CO2-constrained ventilationeconomizer logicenergy efficiency

This study addresses the persistent performance gap between simulation and real-world deployment of reinforcement learning (RL) in industrial energy systems, using a district heating network as a case study. The control task is formulated as a Markov decision process, and key challenges—including partial observability, action space design, reward function shaping, and sim-to-real transfer—are systematically analyzed. For the first time, this work provides a comprehensive empirical investigation of RL deployment barriers in an actual industrial setting, uncovering the root causes of performance discrepancies. Although the system achieved stable operation in practice, its measured performance fell significantly short of simulation results, thereby validating critical bottlenecks in real-world RL applications and establishing an empirical foundation for future co-optimization of algorithms and engineering design.

industrial energy systemspartial observabilityreinforcement learning

Comparative Field Deployment of Reinforcement Learning and Model Predictive Control for Residential HVAC

Oct 01, 2025
OB
Ozan Baris Mulayim
🏛️ Carnegie Mellon University | Purdue University | Trane Technologies | Bosch Center for Artificial Intelligence

This study addresses critical challenges—scalability, safety, interpretability, and sample efficiency—in deploying reinforcement learning (RL) and model predictive control (MPC) for residential HVAC systems. For the first time, it conducts a one-month closed-loop comparative experiment in a real residential setting. A physics-informed MPC and a model-based RL (MBRL) approach, leveraging learned dynamic system models, are co-deployed with a heat pump to jointly optimize energy efficiency and thermal comfort. Results show RL achieves 22% energy savings—marginally exceeding MPC’s 20%—but MPC delivers superior energy efficiency at equivalent comfort levels. The study quantifies RL’s advantage in reducing modeling effort while empirically identifying, for the first time, key practical bottlenecks: unsafe policy initialization, instability during online adaptation, and actuation deviation. These findings establish a crucial benchmark and provide actionable design insights for operationalizing intelligent building control algorithms.

Addressing safety, interpretability, and sample efficiency issues in real-world RL deploymentComparing Reinforcement Learning and Model Predictive Control for residential HVAC systemsEvaluating practical scalability challenges of advanced HVAC control strategies

Latest Papers

What's happening recently
View more

This study addresses the issue of excessive compressor cycling in residential heat pumps, which accelerates equipment wear—a factor commonly overlooked by existing reinforcement learning controllers that focus primarily on energy consumption and thermal comfort. To bridge this gap, the work explicitly incorporates compressor wear into the reward function and evaluates Soft Actor-Critic (SAC) and Proximal Policy Optimization (PPO) algorithms within the hydronic heat pump case of the BOPTEST platform. Results demonstrate that SAC autonomously learns a variable-speed continuous modulation strategy, achieving zero start-stop cycles while reducing thermal discomfort by 90.7% at the cost of only an 11.5% increase in operational expenditure, substantially outperforming conventional control approaches.

compressor wearheat pump controlon-off cycling

This study addresses the challenge of achieving efficient, safe, and adaptive control of water-cooled chillers and air-side systems in tropical commercial buildings. The authors propose the CQD-ERL controller, which introduces, for the first time in HVAC control, a context-aware quality-diversity (QD) evolutionary mechanism. By integrating evolutionary algorithms with the soft actor-critic (SAC) policy gradient method, the approach maintains a policy archive indexed jointly by operational context and behavioral descriptors, enabling the co-optimization of diverse, high-performing control policies. This framework supports dynamic selection of specialized policies tailored to prevailing weather conditions and load profiles. A deterministic safety shielding mechanism is incorporated to enforce critical constraints such as humidity levels and cooling tower approach temperature. Evaluated through full-year backtesting on a commercial building in Singapore, the proposed controller significantly outperforms the ASHRAE Guideline 36 baseline, achieving energy-efficient operation while rigorously maintaining safety constraints.

evolutionary reinforcement learningHVAC controloperating context

This work addresses the limited generalization of reinforcement learning (RL) policies in real-world deployment, where environmental dynamics shift, action and observation spaces vary, and control objectives change. To tackle this challenge, the authors introduce the first large-scale, physically realistic continuous control benchmark for HVAC control, built upon EnergyPlus. Leveraging a parameterized building generator, the benchmark systematically produces diverse building configurations and defines standardized tasks to evaluate key generalization capabilities—including objective adaptation, dynamic shifts, action space variations, and cross-domain transfer. The platform supports heterogeneous observation and action spaces and integrates with the Gymnasium interface alongside a unified evaluation protocol, thereby providing a robust foundation for advancing both building energy efficiency and the robustness of RL algorithms.

benchmarkgeneralizationreal-world deployment

Hot Scholars

MB

Mona Bielig

Konstanz University, Seeburg Castle University
VG

Vikas Garg

Massachusetts Institute of Technology (MIT)
Machine Learning
CK

Celina Kacperski

Konstanz University, Seeburg Castle University
sustainabilityenvironmenttechnologygames
CP

Charalampos P. Andriotis

Assistant Professor, Delft University of Technology
Decision MakingArtificial IntelligenceRisk & ReliabilityOptimization
ZB

Zaharah Bukhsh

Assistant professor, Eindhoven University of Technology
decision-makingpredictive maintenancerepresentation learningdeep reinforcement learning