TRUAV: Distributed Multi-Agent Reinforcement Learning for Trajectory Planning and Routing Enhancement in UAV-Aided IoT-Enabled VANETs

📅 2026-07-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the practical limitations of conventional centralized UAV trajectory planning in dense urban environments, where bandwidth constraints and energy consumption hinder scalability. To overcome these challenges, the authors propose TRUAV, a distributed multi-agent reinforcement learning framework based on independent tabular Q-learning. Each UAV optimizes its trajectory and routing decisions using only local observations—such as vehicle density, queue status, and neighboring UAV positions—without requiring global communication. A potential-game-inspired reward mechanism is introduced to encourage spatial diversity and routing-aware positioning while maintaining energy efficiency. Large-scale simulations in an urban setting with 200 mobile vehicles demonstrate that TRUAV achieves coverage and packet delivery performance comparable to centralized deep reinforcement learning approaches, while significantly reducing relay latency and improving energy efficiency.
📝 Abstract
Unmanned aerial vehicles (UAVs) have emerged as a key enabler of next-generation Internet of Things (IoT) ecosystems, offering flexible aerial relaying to extend connectivity across dynamic vehicular ad hoc networks (VANETs) in smart city environments. However, conventional centralized approaches for UAV trajectory planning require continuous global network state aggregation, making them impractical under bandwidth and energy constraints typical of dense urban deployments. In this article, we present TRUAV, a distributed multi-agent reinforcement learning framework based on independent tabular Q-learning for joint UAV trajectory planning and routing enhancement in UAV-aided VANETs. Each UAV is equipped with a local Q-learning agent that operates purely on locally observable information, including vehicle density, packet queue states, and neighbor UAV positions, thereby eliminating the need for global state exchange. A potential-game-inspired reward design encourages spatial diversity and routing-aware UAV positioning among interacting agents while accounting for energy consumption. Numerical simulations over a large urban area with 200 mobile vehicles show that the proposed TRUAV framework achieves network coverage and packet delivery ratios comparable to centralized deep reinforcement learning methods, while also improving relay delay and energy efficiency. Finally, we discuss emerging challenges and future research directions for distributed multi-agent UAV-assisted IoT systems.
Problem

Research questions and friction points this paper is trying to address.

UAV trajectory planning
VANETs
distributed multi-agent reinforcement learning
IoT-enabled networks
routing enhancement
Innovation

Methods, ideas, or system contributions that make the work stand out.

Distributed Multi-Agent Reinforcement Learning
UAV Trajectory Planning
VANET Routing
Potential-Game Reward Design
Local Observability
🔎 Similar Papers
No similar papers found.