🤖 AI Summary
This study addresses the instability of shared structures in federated LoRA, where parameter similarity is susceptible to initialization interference and neglects local input distributions. To overcome this, we propose FedSAIL, a framework that exploits second-order moment statistics of input-aware action matrices to mine robust shared subspace geometry across clients. Our analysis reveals that the principal singular directions of action matrices are more robust than direct parameter alignment. Accordingly, we establish a novel aggregation paradigm that replaces conventional weight averaging with subspace-guided aggregation, complemented by regularization constraints to optimize low-rank adaptation training. Extensive experiments demonstrate that FedSAIL significantly improves predictive performance across multiple benchmarks while substantially reducing communication overhead.
📝 Abstract
Federated low-rank adaptation (LoRA) requires identifying an update structure that is shared across heterogeneous clients. Prior work reports strong similarity among trained LoRA projection matrices across clients; however, such agreement may be largely induced by common initialization and collapses toward random overlap under independent initialization. More crucially, relying solely on parameter similarity inherently ignores the influence of local input regime. To uncover a more robust shared structure, we introduce an input-aware action matrix that weights the adapter update by the second-moment statistics of local layer inputs. Empirically, while parameter similarity vanishes, the leading right singular directions of this action matrix remain strongly aligned across clients. This shared geometry preserves task-conditioned differences and naturally varies across network depths. Motivated by these findings, we propose Federated Subspace-Guided Action-Informed Learning (FedSAIL). Instead of averaging weights, FedSAIL estimates a shared action subspace to regularize local training while preserving client-specific coefficients. Across several benchmarks, our approach consistently improves predictive performance over competing federated LoRA methods while reducing communication cost significantly.