Tactical Decision for Multi-UGV Confrontation with a Vision-Language Model-Based Commander

📅 2025-07-15
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Autonomous evolution from situational awareness to tactical decision-making in multi-unmanned ground vehicle (UGV) adversarial scenarios remains challenging due to the lack of interpretable, end-to-end cognitive frameworks. Method: This paper proposes a command architecture synergizing a vision-language model (VLM) and a lightweight large language model (LLM), enabling end-to-end coupling of perception understanding and strategic reasoning within a unified semantic space. It introduces the first explainable multi-agent tactical cognition model spanning the full “perception–cognition–decision” pipeline, emulating human commander reasoning. Contribution/Results: Unlike rule-based or black-box reinforcement learning baselines, our approach significantly improves generalizability and decision interpretability. Simulation results demonstrate a win rate exceeding 80%; ablation studies validate the efficacy and robustness of each module.

Technology Category

Multiagent Systems: Adversarial AgentsCognitive Modeling & Cognitive Systems: Agent ArchitecturesMachine Learning: Large Multimodal Models (LMMs)

Application Category

Semantics and Knowledge: Data modeling to support human-machine intelligence, including LLMs agents, intelligent system behavior, explanations, and user-friendly interactionsSearch and Retrieval-Augmented AI: Agentic searchResponsible Web: Machine-in-the-loop, human agency and autonomy
📝 Abstract
In multiple unmanned ground vehicle confrontations, autonomously evolving multi-agent tactical decisions from situational awareness remain a significant challenge. Traditional handcraft rule-based methods become vulnerable in the complicated and transient battlefield environment, and current reinforcement learning methods mainly focus on action manipulation instead of strategic decisions due to lack of interpretability. Here, we propose a vision-language model-based commander to address the issue of intelligent perception-to-decision reasoning in autonomous confrontations. Our method integrates a vision language model for scene understanding and a lightweight large language model for strategic reasoning, achieving unified perception and decision within a shared semantic space, with strong adaptability and interpretability. Unlike rule-based search and reinforcement learning methods, the combination of the two modules establishes a full-chain process, reflecting the cognitive process of human commanders. Simulation and ablation experiments validate that the proposed approach achieves a win rate of over 80% compared with baseline models.
Problem

Research questions and friction points this paper is trying to address.

Autonomous tactical decisions in multi-UGV confrontations
Overcoming limitations of rule-based and reinforcement learning methods
Intelligent perception-to-decision reasoning with interpretability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-language model for scene understanding
Lightweight LLM for strategic reasoning
Shared semantic space unifies perception and decision
🔎 Similar Papers
No similar papers found.
L
Li Wang
Yangtze Delta Region Academy, Beijing Institute of Technology, Jiaxing, 314001, China.
Qizhen Wu
Qizhen Wu
Beihang University
L
Lei Chen
Advanced Research Institute of Multidisciplinary Sciences, Beijing Institute of Technology, Beijing, 100081, China.