Building2Building: A Large Scale Benchmark for Generalizable Real-World Reinforcement Learning

📅 2026-07-17
📈 Citations: 0
✹ Influential: 0
📄 PDF
đŸ€– AI Summary
This work addresses the limited generalization of reinforcement learning (RL) policies in real-world deployment, where environmental dynamics shift, action and observation spaces vary, and control objectives change. To tackle this challenge, the authors introduce the first large-scale, physically realistic continuous control benchmark for HVAC control, built upon EnergyPlus. Leveraging a parameterized building generator, the benchmark systematically produces diverse building configurations and defines standardized tasks to evaluate key generalization capabilities—including objective adaptation, dynamic shifts, action space variations, and cross-domain transfer. The platform supports heterogeneous observation and action spaces and integrates with the Gymnasium interface alongside a unified evaluation protocol, thereby providing a robust foundation for advancing both building energy efficiency and the robustness of RL algorithms.
📝 Abstract
Reinforcement learning (RL) has achieved strong results in control, yet learned policies remain brittle to changes in dynamics, action spaces, observation spaces, or goals, a critical limitation for real-world deployment. Existing benchmarks offer limited diversity and complexity, making it difficult to rigorously study transfer, multi-task learning, and meta-learning in RL. We introduce Building2Building (B2B), a large-scale suite of realistic Heating, Ventilation, and Air Conditioning (HVAC) control environments built on EnergyPlus, a state-of-the-art building simulator. B2B is fully compatible with the Gymnasium interface and features a parametric building generator, enabling the systematic generation of diverse building configurations with heterogeneous observation and action spaces. Based on this suite, we define benchmark tasks targeting key open challenges in RL, including goal adaptation, dynamics adaptation, action-space shifts, and cross-domain transfer. By providing a large-scale, diverse, and physically grounded testbed with standardized evaluation protocols, B2B enables systematic investigation of generalization and transfer in continuous control. Beyond advancing research on generalization in RL, this new benchmark also carries significant societal implications by enabling improved HVAC control at scale, one of the most energy-intensive systems in buildings.
Problem

Research questions and friction points this paper is trying to address.

reinforcement learning
generalization
transfer learning
benchmark
real-world deployment
Innovation

Methods, ideas, or system contributions that make the work stand out.

generalizable reinforcement learning
HVAC control
parametric building generator
cross-domain transfer
EnergyPlus simulation
V
Vincent Taboga
Mila - Quebec AI Institute; UniversitĂ© de MontrĂ©al - DĂ©partement d’Informatique et de Recherche OpĂ©rationnelle
J
Justin Veilleux
Mila - Quebec AI Institute; UniversitĂ© de MontrĂ©al - DĂ©partement d’Informatique et de Recherche OpĂ©rationnelle
D
Doseok Jang
Mila - Quebec AI Institute; UniversitĂ© de MontrĂ©al - DĂ©partement d’Informatique et de Recherche OpĂ©rationnelle
A
Anushree Rankawat
UniversitĂ© de MontrĂ©al - DĂ©partement d’Informatique et de Recherche OpĂ©rationnelle
Pierre-Luc Bacon
Pierre-Luc Bacon
University of Montreal
reinforcement learningartificial intelligence