🤖 AI Summary
Federated learning for real-world multi-domain text data faces dual non-IID challenges—both input (language/semantics) and output (label distribution) exhibit significant heterogeneity across clients. Method: We propose Adaptive Federated Distillation (AFD), the first framework to address this dual non-IID setting. AFD introduces a fine-grained, multilingual non-IID text benchmark and designs a dynamic knowledge transfer mechanism grounded in pretrained language models, integrating federated distillation, gradient alignment, and client-adaptive weight adjustment—all while preserving data privacy. Contribution/Results: Extensive experiments demonstrate that AFD consistently outperforms state-of-the-art methods across diverse multi-domain non-IID text tasks, significantly enhancing global model generalization under both homogeneous and heterogeneous client settings. It effectively captures local data diversity and mitigates performance degradation induced by dual input-output heterogeneity.
📝 Abstract
The widespread success of pre-trained language models has established a new training paradigm, where a global PLM is fine-tuned using task-specific data from local clients. The local data are highly different from each other and can not capture the global distribution of the whole data in real world. To address the challenges of non-IID data in real environments, privacy-preserving federated distillation has been proposed and highly investigated. However, previous experimental non-IID scenarios are primarily identified with the label (output) diversity, without considering the diversity of language domains (input) that is crucial in natural language processing. In this paper, we introduce a comprehensive set of multi-domain non-IID scenarios and propose a unified benchmarking framework that includes diverse data. The benchmark can be used to evaluate the federated learning framework in a real environment. To this end, we propose an Adaptive Federated Distillation (AdaFD) framework designed to address multi-domain non-IID challenges in both homogeneous and heterogeneous settings. Experimental results demonstrate that our models capture the diversity of local clients and achieve better performance compared to the existing works. The code for this paper is available at: https://github.com/jiahaoxiao1228/AdaFD.