Learning from Unreliable Trajectories: Adversarially-Robust Federated Q-Learning

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the degradation of collaborative efficiency caused by adversarial agents in federated reinforcement learning by proposing the Robust Async-Fed-Q algorithm. Built upon a round-asynchronous framework, this method integrates Bellman operator variance reduction with robust server aggregation to preserve statistical gains even in the presence of malicious data. Theoretically, it establishes the first matching information-theoretic lower bound, demonstrating that adversarial effects vanish asymptotically as sample sizes increase. Empirically, the algorithm achieves high-probability finite-time guarantees while substantially reducing communication overhead associated with asynchronous sampling. By tolerating arbitrarily corrupted messages and optimizing overall communication complexity, this work provides a principled and practical solution for resilient federated reinforcement learning under adversarial conditions.
📝 Abstract
We study federated reinforcement learning in which multiple agents interact with a common Markov decision process and communicate through a central server to collaboratively learn the optimal state-action value function. Our goal is to understand whether the sample-efficiency benefits of collaboration can be retained when a fraction of the agents behave adversarially and transmit arbitrarily corrupted information. To address this problem, we introduce Robust Async-Fed-Q, an epoch-based federated learning algorithm that combines variance-reduced estimation of the Bellman optimality operator at the agents with robust aggregation at the server. We establish high-probability finite-time guarantees showing that the proposed method preserves the statistical gains of collaboration among the honest agents while tolerating adversarial corruption. In particular, the effect of the adversarial agents decreases as the amount of data collected by each honest agent grows and eventually vanishes in the infinite-sample limit. We complement these guarantees with information-theoretic lower bounds that characterize the unavoidable statistical cost of adversarial corruption, leading to the first nearly matching upper and lower bounds for adversarially robust federated reinforcement learning. We further extend our framework to accommodate single-trajectory Markovian sampling and heterogeneous partial coverage, where different agents may explore different regions of the state-action space and learning relies on their collective coverage. Finally, our epoch-based design substantially improves the best known communication complexity for federated Q-learning under asynchronous sampling.
Problem

Research questions and friction points this paper is trying to address.

Federated Reinforcement Learning
Adversarial Robustness
Q-Learning
Byzantine Agents
Sample Efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Federated Reinforcement Learning
Adversarial Robustness
Q-Learning
Robust Aggregation
Communication Complexity
🔎 Similar Papers
No similar papers found.