Reinforcement Learning-based Approach for Vehicle-to-Building Charging with Heterogeneous Agents and Long Term Rewards

📅 2025-02-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of coordinated optimization among heterogeneous electric vehicles (EVs) and building energy systems in office park vehicle-to-grid (V2G) scenarios. We propose a reinforcement learning framework integrating Deep Deterministic Policy Gradient (DDPG), action masking, and mixed-integer linear programming (MILP)-guided policy initialization. The method jointly optimizes EV charging/discharging schedules under dynamic and uncertain conditions, simultaneously satisfying user charging requirements, minimizing monthly time-of-use electricity costs, and curtailing net demand peaks. It supports heterogeneous multi-agent coordination, long-horizon optimization with sparse rewards, and generalizable decision-making in continuous action spaces. Trained on real-world EV operational data from an automotive manufacturer, our approach achieves significant electricity cost savings over state-of-the-art baselines and heuristic methods, guarantees 100% user charging satisfaction, and demonstrates strong cross-scenario scalability and engineering deployability.

Technology Category

Planning, Routing, and Scheduling: Learning for Planning and SchedulingSearch and Optimization: Learning to SearchIntelligent Robots: Learning & Optimization for ROB

Application Category

Economics, Online Markets and Human Computation: Cost models of using LLMs in production systemsSystems and Infrastructure for Web, Mobile and WoT: Energy management for devices in mobile Web and WoT environmentsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
📝 Abstract
Strategic aggregation of electric vehicle batteries as energy reservoirs can optimize power grid demand, benefiting smart and connected communities, especially large office buildings that offer workplace charging. This involves optimizing charging and discharging to reduce peak energy costs and net peak demand, monitored over extended periods (e.g., a month), which involves making sequential decisions under uncertainty and delayed and sparse rewards, a continuous action space, and the complexity of ensuring generalization across diverse conditions. Existing algorithmic approaches, e.g., heuristic-based strategies, fall short in addressing real-time decision-making under dynamic conditions, and traditional reinforcement learning (RL) models struggle with large state-action spaces, multi-agent settings, and the need for long-term reward optimization. To address these challenges, we introduce a novel RL framework that combines the Deep Deterministic Policy Gradient approach (DDPG) with action masking and efficient MILP-driven policy guidance. Our approach balances the exploration of continuous action spaces to meet user charging demands. Using real-world data from a major electric vehicle manufacturer, we show that our approach comprehensively outperforms many well-established baselines and several scalable heuristic approaches, achieving significant cost savings while meeting all charging requirements. Our results show that the proposed approach is one of the first scalable and general approaches to solving the V2B energy management challenge.
Problem

Research questions and friction points this paper is trying to address.

Optimizes vehicle-to-building energy management.
Addresses multi-agent reinforcement learning challenges.
Reduces peak energy costs effectively.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reinforcement Learning-based framework
Deep Deterministic Policy Gradient
Action masking and MILP guidance
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
F
Fangqi Liu
Vanderbilt University, Nashville, TN, USA
R
Rishav Sen
Vanderbilt University, Nashville, TN, USA
J
Jose Paolo Talusan
Vanderbilt University, Nashville, TN, USA
A
Ava Pettet
Nissan Advanced Technology Center - Silicon Valley, Santa Clara, CA, USA
A
Aaron Kandel
Nissan Advanced Technology Center - Silicon Valley, Santa Clara, CA, USA
Y
Yoshinori Suzue
Nissan Advanced Technology Center - Silicon Valley, Santa Clara, CA, USA
A
Ayan Mukhopadhyay
Vanderbilt University, Nashville, TN, USA
Abhishek Dubey
Abhishek Dubey
Vanderbilt University
AI Decision ProceduresCyber Physical SystemsPublic TransitEnergy Systems