Score
Design and implement federated aggregation methods that use large-language-model outputs or other semantic/knowledge signals to guide how client updates, embeddings, or model representations are combined. This includes building weighting, selection, and alignment functions that selectively aggregate structural or semantic embeddings, align heterogeneous client representations, and mitigate non‑IID effects while respecting federated constraints (no direct access to raw client data).
Federated learning (FL) confronts fundamental challenges including statistical heterogeneity, privacy preservation, and collaborative efficiency. To address these, this work establishes a hybrid research framework integrating bibliometric analysis and systematic review. It proposes the first multi-level taxonomy of aggregation techniques, structured along three core dimensions: personalization, optimization, and robustness. Through empirical evaluation, it systematically benchmarks mainstream FL architectures and synchronization strategies under both IID and non-IID data settings. Furthermore, it develops a reproducible benchmarking platform for aggregation methods. The study identifies critical technical bottlenecks and provides theoretically grounded guidance and practical pathways for emerging directions—including privacy-enhancing aggregation, heterogeneity-aware modeling, and robust aggregation. Collectively, this work significantly advances the rigor, reproducibility, and extensibility of systematic FL research.
This work addresses two critical challenges in Mixture-of-Experts (MoE) models under federated learning with heterogeneous data: global routing failure caused by divergent client-specific gating preferences and functional ambiguity arising from semantic inconsistency among experts sharing the same index across clients. To jointly mitigate gating divergence and expert semantic drift, we propose FedAlign-MoE, a novel framework that aligns routing preferences through distributional regularization of gating networks and introduces a selective expert parameter aggregation mechanism guided by semantic consistency metrics. Experimental results demonstrate that our approach significantly outperforms existing methods in non-IID settings, achieving faster convergence and higher accuracy.
In heterogeneous federated learning, disparities in model architectures (e.g., CNNs, RNNs, Transformers) and device computational capabilities impede effective aggregation, leading to slow convergence and poor generalization. To address this, we propose a unified aggregation framework with dynamic architecture adaptation. Our key contributions are: (1) a novel differentiable architecture projection (DAP) mechanism that enables parameter-level alignment across heterogeneous architectures without requiring client-side model standardization or modification; and (2) a synergistic aggregation strategy integrating gradient-aware weight distillation and capability-aware elastic weighting, enhancing both collaboration efficiency and robustness. Extensive experiments on CIFAR-10, CIFAR-100, and LEAF datasets demonstrate that our method achieves up to 23.30% higher accuracy than FlexiFed, reduces communication overhead by 37%, and accelerates convergence by 2.1×.
Data heterogeneity—particularly label distribution skew—severely degrades model performance in federated learning. To address this, we propose FedSC, a semantic-aware collaborative framework that introduces, for the first time, a semantic-level prototype collaboration mechanism: (i) relational prototypes to model inter-class semantic structure, and (ii) consensus prototypes to align locally learned knowledge across clients. FedSC further integrates inter-class contrastive learning with divergence-aware aggregation regularization to enhance generalization and convergence stability. We provide theoretical convergence guarantees and empirically demonstrate significant improvements over state-of-the-art methods across diverse highly heterogeneous benchmarks. Ablation studies validate both the effectiveness of semantic prototype collaboration and the necessity of each designed component.
Large language models (LLMs) face significant challenges in multi-model ensembling, including excessive memory overhead, inflexible weight merging, severe knowledge interference, and degraded task performance. Method: This paper proposes an adaptive multi-LLM ensemble framework featuring a novel feedback-regulated selection network and a dynamic weighted fusion mechanism, enabling task-aware dynamic selection of heterogeneous models and knowledge aggregation; it further introduces a feedback-driven loss function and a collaborative distillation strategy across multiple LLMs to mitigate knowledge conflicts. Contribution/Results: Experiments demonstrate that the method reduces knowledge interference by 50%, substantially improving ensemble stability, scalability, and multi-task robustness. Crucially, it achieves these gains without full-parameter fine-tuning, effectively circumventing the memory and adaptability bottlenecks inherent in conventional ensembling approaches.
This work addresses the critical challenge of selecting effective aggregation strategies in federated learning, which significantly impacts model performance yet depends intricately on data heterogeneity, device diversity, and resource constraints, with no existing automated solution. The paper proposes the first end-to-end adaptive framework that seamlessly integrates large language model–based reasoning with lightweight genetic search to automatically identify optimal aggregation strategies without human intervention. Evaluated across diverse non-IID settings and datasets, the approach substantially enhances model robustness and generalization while markedly reducing hyperparameter tuning overhead. This advancement improves the universality and deployment efficiency of federated learning systems, offering a practical pathway toward autonomous, scalable federated optimization.
To address slow convergence and semantic distortion in federated learning (FL) under non-IID data, this paper proposes a semantics-constrained federated aggregation framework. It is the first to embed industrial ontologies (ISA-95/MASON) into distributed optimization, enabling a knowledge-graph-driven semantic regularization and constraint validation mechanism. Theoretically, we derive the first convergence bound for constrained FL: $O(1/sqrt{T} + ho)$, revealing that constraints reduce data heterogeneity by 41% and quantifying the relationship between constraint violation rate $ ho$ and performance collapse thresholds. Empirically, on Bosch manufacturing data (1.18M samples), convergence accelerates by 22% and model divergence decreases by 41.3%. When $ ho < 0.05$, the framework retains 90% of optimal performance; under differential privacy ($varepsilon = 10$), utility loss is only 3.7%, improving the privacy–utility trade-off by 2.7× over standard FL.
This work addresses the limitations of traditional federated fine-tuning—such as high communication overhead, strict model architecture alignment, and reliance on white-box access—which hinder its applicability to large language models and heterogeneous deployment settings. The authors propose a novel semantic consensus–based federated fine-tuning paradigm: clients fine-tune local models on private data and exchange generated text on a shared public prompt set; the server maps these outputs into a semantic space, constructs per-prompt semantic consensus representations, and returns pseudo-labels for local retraining. By eliminating parameter aggregation and transmitting only lightweight generative behaviors, this approach decouples communication complexity from model scale and inherently supports heterogeneous architectures and open-ended generation. Experiments demonstrate that it matches strong baselines in performance while reducing communication costs by up to 1,006× (e.g., with Llama3.1-405B), substantially cutting runtime and energy consumption.
This work addresses the challenge of deploying large foundation models on resource-constrained clients in federated learning by proposing FedSLM, a novel framework that constructs self-contained lightweight client models via SVD-based low-rank decomposition. FedSLM introduces a two-stage aggregation protocol: intra-group synchronization of low-rank adapters followed by inter-group fusion of full-rank representations. The method innovatively incorporates a structure-aligned fusion mechanism operating on nested subspace manifolds, complemented by a confidence-guided auxiliary loss and a weak-to-strong knowledge distillation strategy to enable efficient knowledge transfer. Experimental results demonstrate that FedSLM significantly outperforms existing federated approaches on both natural language and vision-language tasks, maintaining strong representation capabilities under both IID and non-IID data settings while reducing client memory consumption to approximately 50% of that required by the full model.
This work proposes FedOUI, a novel federated learning aggregation strategy that addresses the limitation of existing methods—which predominantly rely on data volume or gradient information while neglecting the internal structural organization of client models in input space. FedOUI introduces, for the first time, an unlabeled, lightweight activation-based metric called the Overfitting-Underfitting Index (OUI) into federated aggregation. By employing a fixed probe batch to dynamically assess structural consistency across client models, and integrating round-level OUI distribution estimation with a smoothing-based reweighting mechanism, FedOUI automatically downweights clients exhibiting anomalous structural behavior. The method significantly enhances both robustness and interpretability of model aggregation, consistently outperforming baselines such as FedAvg and FedProx under strong non-IID and noisy conditions on CIFAR-10, with particularly pronounced gains in highly heterogeneous settings.
This work addresses the challenge of conflicting client updates in federated learning caused by heterogeneous data distributions. The authors propose a geometric projection-based constrained optimization framework for model aggregation, which formulates the global update as a closed-form solution that satisfies layer-wise conflict-free alignment constraints while remaining closest to a reference direction—enabling efficient, non-iterative aggregation. Theoretical analysis establishes guarantees for a common descent structure from the perspective of projection geometry. Experimental results demonstrate that the proposed method significantly improves global model accuracy on standard heterogeneous benchmarks and effectively reduces performance disparities among clients, outperforming current state-of-the-art approaches.