🤖 AI Summary
This study addresses the real-time collaborative scheduling challenge of shared UAV fleets performing on-demand delivery and road network monitoring. We propose an order-based abstraction for monitoring tasks, transforming highly congested nodes into virtual monitoring orders that are pooled with delivery orders to construct a heterogeneous task set. Building upon this formulation, we employ Graph Multi-Agent Q-Learning (Graph-MAQL) combined with dynamic heterogeneous bipartite graph matching to achieve efficient route optimization. The proposed method enables strong operational synergy between delivery and monitoring while supporting zero-shot transferability. Experimental results demonstrate a 25.1% improvement in monitoring performance, a 20.8% enhancement in overall objective optimization, and a reduction in timeout violation rates exceeding 40%.
📝 Abstract
This paper investigates the real-time dispatch of a shared drone fleet for on-demand food delivery and urban road network monitoring. We consider a courier-drone collaborative setting in which couriers transport orders to launchpads and drones complete the final delivery leg to kiosks. Drones may consolidate multiple origin-destination orders within one flight and make monitoring-aware route adjustments to collect real-time traffic information subject to delivery-time constraints. This yields a joint decision problem coupling dynamic order-to-drone matching, multi-order pooling, routing, and time-varying monitoring under fleet-level competition and uncertainty. We propose monitoring-task orderization, which periodically converts road-network nodes with high congestion and stale information into virtual monitoring orders. Pooling these virtual tasks with food-delivery orders creates a unified heterogeneous task set and transforms the coupled matching-and-routing problem into an order-level decision process. Building on this abstraction, we formulate a decentralized graph-interdependent Multi-Agent Markov Decision Process and develop Graph Multi-Agent Q-Learning (Graph-MAQL), which captures localized inter-agent dependencies through bipartite match coordination graphs. Agent-task value estimates are then used as edge weights in a dynamic heterogeneous bipartite matching program for globally feasible execution. Experiments using real-world data reveal strong operational synergy between delivery and monitoring. Monitoring-task orderization improves monitoring performance by 25.1% with less than a 1% reduction in delivery performance, while Graph-MAQL improves the aggregate objective by up to 20.8%, reduces deadline violations by over 40%, and transfers zero-shot to higher demand intensity without retraining.