Online Security Learning in Cooperative Multi-Agent Systems under Hidden Byzantine Attacks

📅 2026-08-06
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of safe learning in online cooperative control under covert Byzantine attacks, where a subset of agents can secretly manipulate joint actions while the learner observes only corrupted executions, thereby compromising system performance. By characterizing how the attacker’s observational capabilities alter the problem’s geometric structure, the authors formulate (s,a)-rectangular and s-rectangular robust Markov decision processes and propose a phase-coupled Robust Estimate-to-Decide (E2D) algorithm. The theoretical analysis establishes, for the first time, an information-theoretic lower bound on safe regret, showing that the expected cumulative deviation gap 𝔼[DK] is unavoidable. The proposed algorithm achieves a near-optimal safe regret bound of Õ(H²S√(AK)) + 𝔼[DK] and proves an Ω(K) lower bound under indistinguishable instances, laying the theoretical foundation for reliable multi-agent learning in Byzantine environments.
📝 Abstract
We study online cooperative control of a multi-agent system under Byzantine attacks. Namely, an unknown, fixed subset of agents are Byzantine comprised and can stealthily overwrite its own coordinates of the team's planned joint action after observing that plan. The learner observes planned actions, public rewards, and public states, but neither the overwrite nor the executed joint action. Our objective is security: to optimize the team performance against the worst overwrites and achieve the optimal security value. We first show that the attacker's information determines the geometry. An attacker that observes the planned action induces an exact $(s,a)$-rectangular robust Markov decision process (MDP) whose rows are convex hulls of overwrite-induced public-outcome laws, whereas a blind attacker induces an $s$-rectangular model. We then identify the information-theoretic limit of security learning, showing that the security regret decomposes exactly into return regret against the response generating the data and a cumulative response gap $D_K$. Two indistinguishable horizon-one instances force $Ω(K)$ expected security regret while return regret is zero, showing that dependence on $D_K$ is unavoidable. Finally, we develop a stage-tied robust estimation-to-decisions learner and prove a regret bound of $\widetilde{\mathcal O}\!\left(H^2S\sqrt{AK}\right)+\mathbb E[D_K]$. Our studies thus provide comprehensive theoretical and algorithmic foundations of reliable multi-agent systems under Byzantine attacks.
Problem

Research questions and friction points this paper is trying to address.

Byzantine attacks
multi-agent systems
online security learning
robust MDP
cooperative control
Innovation

Methods, ideas, or system contributions that make the work stand out.

Byzantine attacks
robust MDP
security regret
multi-agent reinforcement learning
estimation-to-decisions
🔎 Similar Papers
No similar papers found.