🤖 AI Summary
Optimizing both energy efficiency and reliability of liquid-cooling systems in AI-intensive data centers remains challenging. Method: This paper introduces LC-Opt—the first open-source, end-to-end reinforcement learning (RL) benchmark environment for liquid cooling. Built upon a high-fidelity Modelica digital twin and integrated with Gymnasium, it supports centralized and distributed multi-agent RL algorithms. Its novel agent-grid architecture incorporates LLMs to generate natural-language explanations of control policies and actions, while policy distillation yields interpretable decision trees. A thermally coupled waste-heat recovery module further enhances sustainability. Contribution/Results: Experiments demonstrate effective multi-objective optimization—reducing PUE, suppressing hotspot temperatures, and improving dynamic response—thereby significantly increasing operational trustworthiness and deployability. LC-Opt establishes a standardized, scalable research platform for AI-driven intelligent liquid-cooling control.
📝 Abstract
Liquid cooling is critical for thermal management in high-density data centers with the rising AI workloads. However, machine learning-based controllers are essential to unlock greater energy efficiency and reliability, promoting sustainability. We present LC-Opt, a Sustainable Liquid Cooling (LC) benchmark environment, for reinforcement learning (RL) control strategies in energy-efficient liquid cooling of high-performance computing (HPC) systems. Built on the baseline of a high-fidelity digital twin of Oak Ridge National Lab's Frontier Supercomputer cooling system, LC-Opt provides detailed Modelica-based end-to-end models spanning site-level cooling towers to data center cabinets and server blade groups. RL agents optimize critical thermal controls like liquid supply temperature, flow rate, and granular valve actuation at the IT cabinet level, as well as cooling tower (CT) setpoints through a Gymnasium interface, with dynamic changes in workloads. This environment creates a multi-objective real-time optimization challenge balancing local thermal regulation and global energy efficiency, and also supports additional components like a heat recovery unit (HRU). We benchmark centralized and decentralized multi-agent RL approaches, demonstrate policy distillation into decision and regression trees for interpretable control, and explore LLM-based methods that explain control actions in natural language through an agentic mesh architecture designed to foster user trust and simplify system management. LC-Opt democratizes access to detailed, customizable liquid cooling models, enabling the ML community, operators, and vendors to develop sustainable data center liquid cooling control solutions.