🤖 AI Summary
This study addresses the compound heterogeneity problem arising from mixed client data in federated fine-tuning by proposing the FedSEE framework. This approach reveals that the optimal model constitutes a convex combination of task-specific experts, achieving specialized adaptation through frozen representation learning, public-sample-based expert recovery with alignment routing, and an input-dependent dynamic routing mechanism. Experimental results demonstrate that FedSEE effectively mitigates negative transfer, yielding an overall performance improvement of 2.9 points and a 3.7-point gain for the worst-performing quartile. These findings indicate that the proposed framework substantially enhances the performance of long-tail clients while maintaining robust generalization across heterogeneous federated environments.
📝 Abstract
In collaborative foundation model fine-tuning, client data is rarely homogeneous. Instead, clients typically possess unknown mixtures of distinct data distributions, or tasks. Conventional federated learning primarily addresses heterogeneity across clients without explicitly resolving latent task mixtures within each client. We study this setting as compound heterogeneity, where data is heterogeneous both across and within clients. We study adaptation over a common frozen representation and show that, when tasks share the same feature geometry, the optimal model for a client's task mixture under squared loss is a convex combination of the optimal models for its underlying tasks. Thus, a single locally trained model represents the client's overall task mixture, while individual inputs may be drawn from different underlying task distributions. This motivates routing inputs to specialized experts, and we show that, when the task optima form a simplex, task-aligned routing achieves lower risk than any single adapted model for genuinely mixed clients. With access to a small set of task-labeled public samples, we derive a convex program to recover task experts and match them to their corresponding tasks. Our routing analysis shows that effective specialization requires input-dependent expert selection aligned with each client's task mixture. Motivated by this analysis, we propose FedSEE. Across our experiments, FedSEE avoids the negative transfer observed in the evaluated baselines and improves performance by 2.9 points overall and 3.7 points for the worst-served quartile.