🤖 AI Summary
This study addresses the instability and policy collapse in existing bus dispatching methods under stochastic traffic and passenger demand, which arise from conflating irreducible aleatoric uncertainty with epistemic uncertainty due to limited data. To resolve this, the authors propose RE-SAC, a novel framework that explicitly decouples these two uncertainty sources for the first time in bus control. Specifically, an IPM-regularized critic network mitigates aleatoric risk, while a diverse ensemble of Q-functions alleviates overconfident value estimates in data-scarce regions. The proposed robust Bellman operator—free of inner-loop perturbations—provides a theoretical lower bound that reveals how ensemble variance misattributes estimation noise to data gaps. Experiments on a bidirectional bus corridor demonstrate that RE-SAC achieves a cumulative reward of −0.4×10⁶, outperforming SAC (−0.55×10⁶), and reduces MAE in Oracle Q-value estimation under out-of-distribution rare states by 62%, from 4343 to 1647.
📝 Abstract
Bus holding control is challenging due to stochastic traffic and passenger demand. While deep reinforcement learning (DRL) shows promise, standard actor-critic algorithms suffer from Q-value instability in volatile environments. A key source of this instability is the conflation of two distinct uncertainties: aleatoric uncertainty (irreducible noise) and epistemic uncertainty (data insufficiency). Treating these as a single risk leads to value underestimation in noisy states, causing catastrophic policy collapse. We propose a robust ensemble soft actor-critic (RE-SAC) framework to explicitly disentangle these uncertainties. RE-SAC applies Integral Probability Metric (IPM)-based weight regularization to the critic network to hedge against aleatoric risk, providing a smooth analytical lower bound for the robust Bellman operator without expensive inner-loop perturbations. To address epistemic risk, a diversified Q-ensemble penalizes overconfident value estimates in sparsely covered regions. This dual mechanism prevents the ensemble variance from misidentifying noise as a data gap, a failure mode identified in our ablation study. Experiments in a realistic bidirectional bus corridor simulation demonstrate that RE-SAC achieves the highest cumulative reward (approx. -0.4e6) compared to vanilla SAC (-0.55e6). Mahalanobis rareness analysis confirms that RE-SAC reduces Oracle Q-value estimation error by up to 62% in rare out-of-distribution states (MAE of 1647 vs. 4343), demonstrating superior robustness under high traffic variability.