Reinforcement Learning for Delivery Drone-Based Participatory Sensing in Dynamic Environments

📅 2026-07-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the scalability bottleneck and heterogeneous multi-timescale decision-making challenges faced by drones performing joint delivery and sensing tasks in dynamic wind-disturbed environments. To tackle these issues, the paper proposes a two-timescale reinforcement learning framework (TSRL) that decouples decision-making into two coordinated levels: macro-level task scheduling and micro-level velocity control. The framework achieves environmental adaptability through task embedding encoding at the macro level and wind-aware velocity scheduling at the micro level. By integrating sequential drone fitness evaluation with hierarchical reinforcement learning, TSRL demonstrates significant performance gains, improving average system rewards by 20.1% and 46.6% on real-world datasets from Hangzhou and Shanghai, respectively, substantially outperforming existing baselines and validating its effectiveness and scalability in dynamic operational settings.
📝 Abstract
Using Unmanned Aerial Vehicle (UAV) for urban sensing has emerged as a powerful paradigm to monitor the status of the city, e.g., air quality and noise levels, through agile aerial crowdsourcing. Despite this potential, existing UAV-based sensing approaches overlook environmental disturbances like wind that drastically impact drone velocity and energy efficiency. Consequently, directly applying existing methods to this joint delivery and sensing paradigm in dynamic environments faces two severe challenges: (1) scalability bottlenecks as fleet sizes expand; and (2) multi-timescale decision heterogeneity between macro task dispatching and micro velocity control. To tackle these, we formalize the problem as SensUAV and propose a Two TimeScale Reinforcement Learning framework (TSRL). Specifically, TSRL separates decision-making into two cooperative layers. At the macro level, a task-embedding sensing dispatcher handles scalability by separately encoding distinct task features and sequentially evaluating UAV suitability before task selection. At the micro level, a wind-aware velocity controller learns fine-grained velocity scheduling to adapt to dynamic environmental variations. Extensive experiments on real-world datasets demonstrate that TSRL significantly outperforms baselines, achieving average system profit improvements of 20.1% in Hangzhou and 46.6% in Shanghai.
Problem

Research questions and friction points this paper is trying to address.

UAV-based sensing
dynamic environments
environmental disturbances
scalability
multi-timescale decision
Innovation

Methods, ideas, or system contributions that make the work stand out.

Two-TimeScale Reinforcement Learning
UAV-based Participatory Sensing
Wind-aware Velocity Control
Task-embedding Dispatching
Dynamic Urban Sensing