🤖 AI Summary
This study addresses the underutilization of exogenous variables and the limited accuracy of long-horizon forecasting in ride-hailing demand prediction. To this end, we construct the first large-scale benchmark dataset integrating multiple exogenous disturbances, including weather conditions and holidays. Through a systematic evaluation of 30 spatiotemporal sequence models and foundational models under exogenous variable integration scenarios, this work reveals the inherent limitations of existing approaches. Our findings demonstrate that while exogenous variables significantly enhance short-term forecasting performance, current models continue to encounter bottlenecks in complex scenarios and long-horizon predictions. Ultimately, this research establishes a critical evaluation benchmark for the field and delineates promising directions for future optimization.
📝 Abstract
We release Ride-Hailing, a large-scale ride-hailing time series dataset synthesized from DiDi's marketplace data across 200 spatial areas. Ride-Hailing spans four consecutive years at half-hourly granularity and covers three representative exogenous scenarios: Weather Disturbance, Holiday Effect, and Large-scale Event Impact. Built upon Ride-Hailing, we introduce RideBench, a comprehensive benchmark for exogenous-aware ride-hailing forecasting, covering both regular week-ahead forecasting and long-horizon 8-week-ahead forecasting with up to 2,688 prediction steps. RideBench evaluates over 30 representative forecasting methods, including endogenous-only models, exogenous-aware models, and time series foundation models. Our results show that future-known exogenous variables provide clear benefits in regular week-ahead forecasting, especially under weather, holiday, and large-scale event (e.g., major sporting events and concerts) scenarios. However, current exogenous-aware models still struggle to fully capture disturbance-induced pattern changes under complex external contexts. For long-horizon forecasting, existing models cannot simultaneously achieve low pointwise errors, accurate broad trends, and reliable near-term forecasts. These findings reveal a clear mismatch between existing forecasting models and real-world ride-hailing requirements, highlighting the need for models that can better exploit future-known exogenous information, scale across heterogeneous areas, and support long-horizon planning. By introducing Ride-Hailing and RideBench, we aim to encourage the community to study these practical challenges in real-world ride-hailing forecasting.