🤖 AI Summary
This study addresses the dynamic discrepancies between offline and online simulations, as well as incomplete passenger trips, in multi-line bus dispatching. To tackle these challenges, this work proposes a cross-fidelity control strategy based on hybrid offline-online reinforcement learning (H2O). Methodologically, a completion-aware mechanism is introduced to effectively mitigate the failure mode arising from conflicts between low generalized passenger time and trip incompleteness. Furthermore, cross-fidelity simulation interaction techniques are incorporated to rectify the simulation bias between offline data and real-world environments. The proposed approach significantly enhances the robustness and generalization capability of bus dispatching strategies in complex, real-world scenarios.
📝 Abstract
Exploratory reinforcement learning (RL) on an operating bus fleet is impractical,while policies trained only from historical data cannot acquire new experience. Hybrid Offline-and-Online (H2O) RL combines fixed target replay with simulator interaction, but the inexpensive online simulator can differ from the target in transition and event-duration dynamics. We study this cross-fidelity problem for multi-line bus holding and address a failure mode in which lower generalized passenger time coexists with incomplete passenger journeys.