🤖 AI Summary
This paper addresses the safe and economic charging scheduling of electric buses under multi-source uncertainties—including photovoltaic generation volatility, dynamic electricity pricing, uncertain driving durations, and limited charging infrastructure.
Method: We propose a constrained bilevel deep reinforcement learning framework: the upper level employs PPO-Lagrangian for charger allocation, while the lower level adopts MAPPO-Lagrangian to optimize spatiotemporal charging power. Battery state-of-charge (SoC) lower-bound hard constraints are handled via Lagrangian relaxation, and an option-based mechanism enables spatiotemporal abstraction and safety-aware exploration. The problem is formulated as a Constrained Markov Decision Process (CMDP) and solved under the centralized training with decentralized execution (CTDE) paradigm.
Results: Experiments on real-world data demonstrate that our method ensures zero battery depletion while significantly reducing operational costs and improving safety compliance. Moreover, it achieves faster convergence compared to existing baseline methods.
📝 Abstract
The integration of Electric Buses (EBs) with renewable energy sources such as photovoltaic (PV) panels is a promising approach to promote sustainable and low-carbon public transportation. However, optimizing EB charging schedules to minimize operational costs while ensuring safe operation without battery depletion remains challenging - especially under real-world conditions, where uncertainties in PV generation, dynamic electricity prices, variable travel times, and limited charging infrastructure must be accounted for. In this paper, we propose a safe Hierarchical Deep Reinforcement Learning (HDRL) framework for solving the EB Charging Scheduling Problem (EBCSP) under multi-source uncertainties. We formulate the problem as a Constrained Markov Decision Process (CMDP) with options to enable temporally abstract decision-making. We develop a novel HDRL algorithm, namely Double Actor-Critic Multi-Agent Proximal Policy Optimization Lagrangian (DAC-MAPPO-Lagrangian), which integrates Lagrangian relaxation into the Double Actor-Critic (DAC) framework. At the high level, we adopt a centralized PPO-Lagrangian algorithm to learn safe charger allocation policies. At the low level, we incorporate MAPPO-Lagrangian to learn decentralized charging power decisions under the Centralized Training and Decentralized Execution (CTDE) paradigm. Extensive experiments with real-world data demonstrate that the proposed approach outperforms existing baselines in both cost minimization and safety compliance, while maintaining fast convergence speed.