🤖 AI Summary
This work proposes a method for inferring pursuer parameters and planning safe, time-optimal paths in a pursuit-evasion scenario with bounded turn rates. By deploying sacrificial agents that execute straight-line trajectories and observing binary outcomes—interception or survival—of their interactions with the adversary, the approach leverages both boundary and interior geometric reachable set models to inversely estimate pursuer dynamics. A custom loss function combined with multi-start gradient-based optimization enables robust parameter recovery, while Bayesian experimental design guided by the D-optimality criterion efficiently selects the most informative sacrificial trajectories. The method accurately identifies pursuer parameters within only 5–12 interactions. Leveraging these estimates, high-value agent trajectories are synthesized to simultaneously guarantee safety—by avoiding all feasible engagement zones—and achieve time optimality.
📝 Abstract
This paper presents a learning-based framework for estimating pursuer parameters in turn-rate-limited pursuit-evasion scenarios using sacrificial agents. Each sacrificial agent follows a straight-line trajectory toward an adversary and reports whether it was intercepted or survived. These binary outcomes are related to the pursuer's parameters through a geometric reachable-region (RR) model. Two formulations are introduced: a boundary-interception case, where capture occurs at the RR boundary, and an interior-interception case, which allows capture anywhere within it. The pursuer's parameters are inferred using a gradient-based multi-start optimization with custom loss functions tailored to each case. Two trajectory-selection strategies are proposed for the sacrificial agents: a geometric heuristic that maximizes the spread of expected interception points, and a Bayesian experimental-design method that maximizes the D-score of the expected Gauss-Newton information matrix, thereby selecting trajectories that yield maximal information gain. Monte Carlo experiments demonstrate accurate parameter recovery with five to twelve sacrificial agents. The learned engagement models are then used to generate safe, time-optimal paths for high-value agents that avoid all feasible pursuer engagement regions.