🤖 AI Summary
This study addresses the challenge of spatiotemporal traffic dynamics in UAV base station relocation, where existing solutions lack explicit demand modeling. We propose a latent-space decision-planning framework built upon a decomposed spatiotemporal world model. Methodologically, we introduce a recurrent state space model that uniquely integrates an exponential moving average (EMA) latent prediction objective with variance regularization, alongside a differentiable service simulator to reconstruct channel correlations within the latent space. The controller achieves swarm coordination through imagination-based rollouts and uncertainty penalization. Experiments on three real-world datasets demonstrate that our method consistently ranks first in service ratio, reaching up to 0.908, significantly outperforming the strongest baselines. These results validate the effectiveness of observation during decision-making.
📝 Abstract
Uncrewed aerial vehicle base stations (UAV-BSs) are expected to cover traffic demand that shifts across space and time, yet most repositioning schemes either re-solve an optimization problem per slot or learn reactive policies without an explicit demand model. We cast demand-driven fleet repositioning as latent-space decision-time planning and propose DSWM, a decomposed spatio-temporal world model: an agentic controller that perceives the demand field through a rolling observation window, retains operational context in a latent recurrent state, reasons about candidate motions by imagined rollouts under an uncertainty penalty, and coordinates the fleet through replanned first actions. DSWM learns a recurrent state-space model shaped by an exponential-moving-average (EMA) based latent predictive objective with variance regularization. It attaches a differentiable service simulator that replays the association, probabilistic line-of-sight channel, and Shannon rate chain inside latent rollouts. Planning uses a cross-entropy method whose imagined demand is anchored on the current observation window with mixing coefficient $ρ=0.95$. On a unified pipeline over three real datasets (Milan CDR (call detail record), Shanghai Telecom, YJMob100K) and 14 methods including five reproduced IEEE baselines, DSWM attains weekday served ratios of 0.889, 0.908, and 0.898, ranking first among non-ablated configurations on every dataset. On Milan it improves over the strongest non-learning baseline (Greedy, 0.780) by 0.109, a margin that comes from decision-time use of observations rather than prediction accuracy.