🤖 AI Summary
This study addresses the challenges of client distribution mismatch and constrained parameter synchronization in heterogeneous federated reinforcement learning by proposing a novel framework that integrates diffusion priors with an optimal transport mixture-of-experts (OT-MoE) mechanism. Methodologically, diffusion models serve as behavioral priors to align heterogeneous policies, while the OT-MoE mechanism enables personalized aggregation within the distribution space, overcoming the limitations of conventional direct averaging. Additionally, a DICE value baseline is introduced to provide low-variance optimization guidance. Experimental results demonstrate that the proposed approach significantly improves average returns and final performance in heterogeneous environments, while effectively enhancing worst-round robustness and overall learning stability.
📝 Abstract
Federated Reinforcement Learning (FRL) enables collaborative policy learning across distributed agents with heterogeneous environments. While recent methods based on variance reduction, divergence penalization, and momentum optimization improve FRL under heterogeneous settings, they still primarily synchronize policy or value-network parameters and do not explicitly address distributional mismatch among heterogeneous clients. Therefore, we propose \textbf{FedGuide}, a FRL framework that uses diffusion priors as behavior models to provide personalized data supported distributions for heterogeneous local policy learning. Instead of directly averaging local policies, FedGuide aggregates those diffusion priors through Optimal-Transport Mixture-of-Experts (OT-MoE), preserving heterogeneous behavior modes in distribution space. It further develops a Distribution Correction Estimation (DICE) value baseline to provide low-variance, return-aware guidance for local policy improvement. Experiments across heterogeneous environments show that FedGuide outperforms representative FRL methods in client-average returns, final-round performance, and worst-round robustness, while maintaining stable learning under stronger heterogeneity.