๐ค AI Summary
This work addresses the challenge of dynamically varying attacker types, strengths, and target sets in repeated Stackelberg security games, where attack sequences are non-stationary and unpredictable. The authors propose an extended game-theoretic model that accommodates multi-target attacks and develop a no-regret online learning algorithm to handle such uncertainty. Under full-information feedback, the approach integrates a mixed-integer linear programming oracle based on an optimistic tie-breaking rule with the Follow-the-Perturbed-Leader (FPL) algorithm; under bandit feedback, it reconstructs utility estimates via barycentric spanner-based techniques. Theoretical analysis establishes sublinear regret bounds of $O(\sqrt{T})$ and $O(T^{2/3})$ in the full-information and bandit settings, respectivelyโthe first such guarantees for time-varying multi-target attack scenarios. Empirical simulations further demonstrate the robustness and effectiveness of the proposed method in dynamic environments.
๐ Abstract
This work studies no-regret online learning in Repeated Stackelberg Security Games with time-varying attack intensities. We formulate an extended security game in which an attacker may select multiple targets and derive an exact mixed-integer linear programming oracle under a optimistic tie-breaking rule. Under full-information feedback, the oracle is integrated with Follow-the-Perturbed-Leader and yields expected $\mathcal{O}(\sqrt{T})$ regret against non-anticipating sequences with time-varying follower numbers, attack intensities, and attacker types. Under bandit feedback, we consider multiple followers sharing a fixed attacker type and use a barycentric-spanner construction to reconstruct utility estimates from aggregate attack observations, obtaining expected $\mathcal{O}(T^{2/3})$ regret. Extensive simulations demonstrate the robustness and effectiveness of our approach under full and partial information feedback.