Bidirectional Task-Motion Planning Based on Hierarchical Reinforcement Learning for Strategic Confrontation

📅 2025-04-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
In swarm robotic strategic adversarial scenarios, conventional unidirectional decoupling between task planning (discrete commands) and motion planning (continuous actions) severely limits environmental adaptability. Method: This paper proposes a bidirectional interactive hierarchical reinforcement learning framework. It innovatively introduces a cross-layer collaborative training mechanism and a task-trajectory mapping prediction model to enable dynamic joint optimization of high- and low-level decisions; further, it incorporates bidirectional policy networks and cross-layer parameter sharing, breaking the traditional task-and-motion planning (TAMP) pipeline paradigm. Contribution/Results: Evaluated in large-scale simulations and real-robot experiments, the system achieves an adversarial win rate exceeding 80% while maintaining sub-10-ms per-step decision latency. It significantly improves generalization capability and real-time performance, establishing a scalable new paradigm for dynamic decision-making in swarm intelligence.

Technology Category

Multiagent Systems: Adversarial AgentsHumans and AI: Human-Aware Planning and Behavior PredictionIntelligent Robots: Learning & Optimization for ROB

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingResponsible Web: Machine-in-the-loop, human agency and autonomyEconomics, Online Markets and Human Computation: Uses of LLMs and GenAI for marketplace design, bidding, and strategic interactions
📝 Abstract
In swarm robotics, confrontation scenarios, including strategic confrontations, require efficient decision-making that integrates discrete commands and continuous actions. Traditional task and motion planning methods separate decision-making into two layers, but their unidirectional structure fails to capture the interdependence between these layers, limiting adaptability in dynamic environments. Here, we propose a novel bidirectional approach based on hierarchical reinforcement learning, enabling dynamic interaction between the layers. This method effectively maps commands to task allocation and actions to path planning, while leveraging cross-training techniques to enhance learning across the hierarchical framework. Furthermore, we introduce a trajectory prediction model that bridges abstract task representations with actionable planning goals. In our experiments, it achieves over 80% in confrontation win rate and under 0.01 seconds in decision time, outperforming existing approaches. Demonstrations through large-scale tests and real-world robot experiments further emphasize the generalization capabilities and practical applicability of our method.
Problem

Research questions and friction points this paper is trying to address.

Integrates discrete commands and continuous actions for swarm robotics confrontation
Bridges task allocation and path planning via bidirectional hierarchical reinforcement learning
Enhances adaptability in dynamic environments through cross-training and trajectory prediction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bidirectional hierarchical reinforcement learning for dynamic interaction
Cross-training enhances hierarchical learning efficiency
Trajectory prediction links tasks to actionable goals
💼 Related Jobs
No related jobs found.
Qizhen Wu
Qizhen Wu
Beihang University
L
Lei Chen
Advanced Research Institute of Multidisciplinary Sciences and State Key Laboratory of CNS/ATM, Beijing Institute of Technology, Beijing 100081, China
K
Kexin Liu
School of Automation Science and Electrical Engineering, Beihang University, Beijing 100191, China
Jinhu Lu
Jinhu Lu
School of Automation Science and Electrical Engineering, Beihang University, Beijing 100191, China