GAMBIT: Learning to Plan Continuous Multi-Robot Trajectories

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of achieving self-sacrificial behaviors and collision-free navigation in multi-robot coordination within dense environments. To this end, it proposes a safe planning framework that integrates imitation learning with reinforcement fine-tuning. Methodologically, drawing inspiration from game-theoretic strategies, a self-sacrificial coordination policy is designed. Continuous motion primitives are generated using a double-integrator dynamics model, while a receding-horizon safety mechanism incorporating backup trajectories ensures collision-free execution. Experimental results demonstrate that the proposed approach significantly outperforms existing baselines, supporting real-time coordination for thousands of robots with end-to-end latency below several hundred milliseconds and exhibiting excellent scalability.
📝 Abstract
GAMBIT is an opening chess move in which a player sacrifices a piece, typically a pawn, to gain a positional advantage later in the game. Analogously, in multi-robot coordination, individual robots may need to forgo locally reward-maximising behaviours to improve overall team performance. Such self-sacrificial behaviours are difficult to capture with manually designed heuristics, particularly in dense, interaction-rich environments. Focusing on double-integrator continuous dynamics, this work studies how to learn such coordinated heuristics over motion primitives for multi-robot trajectory execution. Our framework, GAMBIT, first learns coordinated motion-primitive selection through imitation learning and subsequently fine-tunes the policy through reinforcement learning. We further introduce a safeguarded rollout mechanism with backup trajectories that guarantees collision-free execution at all times. Experiments demonstrate that GAMBIT substantially outperforms a range of baselines, including centralised motion planners and decentralised reactive planners, while exhibiting strong scalability. In particular, it coordinates over a thousand robots with planning latency below a few hundred milliseconds in continuous domains.
Problem

Research questions and friction points this paper is trying to address.

multi-robot trajectory planning
continuous dynamics
coordination
collision avoidance
scalability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Robot Trajectory Planning
Motion Primitives
Imitation Learning
Reinforcement Learning
Collision Avoidance