Optimizing Electric Bus Charging Scheduling With Uncertainties Using Hierarchical Deep Reinforcement Learning

📅 2025-05-15
🏛️ IEEE Internet of Things Journal
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the multi-timescale charging scheduling problem for electric bus fleets under uncertainties in travel time, energy consumption, and electricity pricing, this paper proposes DAC-MAPPO-E, a hierarchical reinforcement learning framework. The method introduces a novel two-level architecture: a top-level attention-based module that decouples global state representations to enable scalable, coordinated decision-making; and a bottom-level decentralized power optimization layer built upon Multi-Agent Proximal Policy Optimization (MAPPO). It integrates Hierarchical Markov Decision Process (H-MDP) modeling, a dual Actor-Critic structure, and real-time IoT data-driven adaptation. Experimental results demonstrate that DAC-MAPPO-E reduces charging costs and grid peak-to-valley load differences compared to baseline methods, supports real-time scheduling for fleets exceeding one thousand vehicles, achieves 40% faster convergence, and reduces computational overhead by 35%. The framework exhibits strong scalability, robustness against uncertainties, and practical engineering applicability.

Technology Category

Planning, Routing, and Scheduling: Planning/Scheduling and LearningSearch and Optimization: Learning to SearchReasoning under Uncertainty: Sequential Decision Making

Application Category

Systems and Infrastructure for Web, Mobile and WoT: Energy management for devices in mobile Web and WoT environmentsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systemsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
📝 Abstract
The growing adoption of Electric Buses (EBs) represents a significant step toward sustainable development. By utilizing Internet of Things (IoT) systems, charging stations can autonomously determine charging schedules based on real-time data. However, optimizing EB charging schedules remains a critical challenge due to uncertainties in travel time, energy consumption, and fluctuating electricity prices. Moreover, to address real-world complexities, charging policies must make decisions efficiently across multiple time scales and remain scalable for large EB fleets. In this paper, we propose a Hierarchical Deep Reinforcement Learning (HDRL) approach that reformulates the original Markov Decision Process (MDP) into two augmented MDPs. To solve these MDPs and enable multi-timescale decision-making, we introduce a novel HDRL algorithm, namely Double Actor-Critic Multi-Agent Proximal Policy Optimization Enhancement (DAC-MAPPO-E). Scalability challenges of the Double Actor-Critic (DAC) algorithm for large-scale EB fleets are addressed through enhancements at both decision levels. At the high level, we redesign the decentralized actor network and integrate an attention mechanism to extract relevant global state information for each EB, decreasing the size of neural networks. At the low level, the Multi-Agent Proximal Policy Optimization (MAPPO) algorithm is incorporated into the DAC framework, enabling decentralized and coordinated charging power decisions, reducing computational complexity and enhancing convergence speed. Extensive experiments with real-world data demonstrate the superior performance and scalability of DAC-MAPPO-E in optimizing EB fleet charging schedules.
Problem

Research questions and friction points this paper is trying to address.

Optimizing EB charging schedules under travel and price uncertainties
Enabling multi-timescale decisions for large-scale EB fleets
Reducing computational complexity in decentralized charging policies
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hierarchical Deep Reinforcement Learning for scheduling
Double Actor-Critic with attention mechanism
Multi-Agent Proximal Policy Optimization enhancement
💼 Related Jobs
No related jobs found.
Jiaju Qi
Jiaju Qi
University of Guelph
L
Lei Lei
School of Engineering, University of Guelph, Guelph, ON N1G 2W1, Canada
T
Thorsteinn Jonsson
EthicalAI, Waterloo, ON N2L 0C7, Canada
D
Dusit Niyato
College of Computing and Data Science, Nanyang Technological University, 50 Nanyang Avenue, Singapore 639798