🤖 AI Summary
To address the multi-timescale charging scheduling problem for electric bus fleets under uncertainties in travel time, energy consumption, and electricity pricing, this paper proposes DAC-MAPPO-E, a hierarchical reinforcement learning framework. The method introduces a novel two-level architecture: a top-level attention-based module that decouples global state representations to enable scalable, coordinated decision-making; and a bottom-level decentralized power optimization layer built upon Multi-Agent Proximal Policy Optimization (MAPPO). It integrates Hierarchical Markov Decision Process (H-MDP) modeling, a dual Actor-Critic structure, and real-time IoT data-driven adaptation. Experimental results demonstrate that DAC-MAPPO-E reduces charging costs and grid peak-to-valley load differences compared to baseline methods, supports real-time scheduling for fleets exceeding one thousand vehicles, achieves 40% faster convergence, and reduces computational overhead by 35%. The framework exhibits strong scalability, robustness against uncertainties, and practical engineering applicability.
📝 Abstract
The growing adoption of Electric Buses (EBs) represents a significant step toward sustainable development. By utilizing Internet of Things (IoT) systems, charging stations can autonomously determine charging schedules based on real-time data. However, optimizing EB charging schedules remains a critical challenge due to uncertainties in travel time, energy consumption, and fluctuating electricity prices. Moreover, to address real-world complexities, charging policies must make decisions efficiently across multiple time scales and remain scalable for large EB fleets. In this paper, we propose a Hierarchical Deep Reinforcement Learning (HDRL) approach that reformulates the original Markov Decision Process (MDP) into two augmented MDPs. To solve these MDPs and enable multi-timescale decision-making, we introduce a novel HDRL algorithm, namely Double Actor-Critic Multi-Agent Proximal Policy Optimization Enhancement (DAC-MAPPO-E). Scalability challenges of the Double Actor-Critic (DAC) algorithm for large-scale EB fleets are addressed through enhancements at both decision levels. At the high level, we redesign the decentralized actor network and integrate an attention mechanism to extract relevant global state information for each EB, decreasing the size of neural networks. At the low level, the Multi-Agent Proximal Policy Optimization (MAPPO) algorithm is incorporated into the DAC framework, enabling decentralized and coordinated charging power decisions, reducing computational complexity and enhancing convergence speed. Extensive experiments with real-world data demonstrate the superior performance and scalability of DAC-MAPPO-E in optimizing EB fleet charging schedules.