SpeechLLM Meets Federated Learning for End-to-End ASR: English and Italian Case Studies

📅 2026-07-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of efficiently training end-to-end automatic speech recognition (ASR) systems based on large language models while preserving data privacy. We present the first systematic exploration of training Speech Large Language Models (SpeechLLMs) under a federated learning framework, proposing optimization strategies tailored to their high-dimensional parameters and inherent communication bottlenecks. Furthermore, we analyze the impact of encoder architectures on model performance. Experimental results demonstrate that our approach achieves competitive word error rates across diverse acoustic conditions and speaking styles in both English and Italian monolingual ASR tasks, while substantially reducing communication overhead. These findings lay a practical foundation for deploying multilingual federated ASR systems in real-world scenarios.
📝 Abstract
Federated learning (FL) enables privacy-preserving training of automatic speech recognition (ASR) systems across distributed data sources, yet its application to large-scale speech language models (SpeechLLMs) remains unexplored. This paper presents the first systematic study of federated training for SpeechLLM-based end-to-end ASR systems. We design a communication-efficient federated optimization strategy tailored to the unique challenges of SpeechLLM architectures, addressing high-dimensional parameter spaces, gradient communication overhead, and computational constraints in distributed settings. Through extensive empirical evaluation on monolingual ASR tasks in English and Italian, we demonstrate the effectiveness and stability of our federated approach compared to centralized training baselines across diverse acoustic conditions and speaking styles. Additionally, we conduct a comprehensive ablation study analyzing the impact of different speech encoder architectures on monolingual English ASR performance within the federated framework, providing insights into optimal model configurations for decentralized training. Our results achieve competitive word error rates while reducing communication costs, establishing practical foundations for federated SpeechLLM deployment in real-world multilingual scenarios.
Problem

Research questions and friction points this paper is trying to address.

Federated Learning
SpeechLLM
Automatic Speech Recognition
Privacy-Preserving Training
End-to-End ASR
Innovation

Methods, ideas, or system contributions that make the work stand out.

Federated Learning
SpeechLLM
End-to-End ASR
Communication-Efficient Optimization
Monolingual Speech Recognition
🔎 Similar Papers
M
Mohamed Nabih Ali
Center of Augmented Intelligence, Fondazione Bruno Kessler, Trento, Italy
D
Daniele Falavigna
Center of Augmented Intelligence, Fondazione Bruno Kessler, Trento, Italy
Alessio Brutti
Alessio Brutti
FBK
audio/speech processing