CAFL-L: Constraint-Aware Federated Learning with Lagrangian Dual Optimization for On-Device Language Models

📅 2025-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the stability and efficiency challenges of federated language model training on resource-constrained edge devices—specifically under joint energy, communication, memory, and thermal limitations—this paper proposes the first personalized federated learning framework that unifies multi-dimensional resource constraints into a single optimization model. Methodologically, it introduces Lagrangian dual optimization to dynamically coordinate layer freezing, local update steps, batch size, and communication compression ratio, while integrating gradient accumulation to respect token budget constraints. Its key innovation lies in the first joint modeling of energy, communication, memory, and thermal constraints, enabling resource-aware adaptive training. Evaluated on character-level language models, the framework achieves a 20% reduction in memory footprint and a 95% decrease in communication overhead compared to FedAvg, while maintaining competitive validation accuracy.

Technology Category

Machine Learning: Distributed Machine Learning & Federated LearningNatural Language Processing: Learning & Optimization for NLPSearch and Optimization: Learning to Search

Application Category

User Modeling, Personalization and Recommendation: Federated recommendation systems and personalizationSystems and Infrastructure for Web, Mobile and WoT: Federated Web and WoT systems, including distributed, federated and edge-based data processingSearch and Retrieval-Augmented AI: Multilingual and cross-lingual Web search
📝 Abstract
We introduce Constraint-Aware Federated Learning with Lagrangian Dual Optimization (CAFL-L), a principled extension of FedAvg that explicitly incorporates device-level resource constraints including energy, communication, memory, and thermal budgets. CAFL-L employs Lagrangian dual optimization to dynamically adapt training hyperparameters -- freezing depth, local steps, batch size, and communication compression -- while preserving training stability through token-budget preservation via gradient accumulation. Experiments on a character-level language model demonstrate that CAFL-L achieves superior constraint satisfaction compared to standard FedAvg (reducing memory usage by 20% and communication by 95%) while maintaining competitive validation performance, making it practical for deployment on resource-constrained edge devices.
Problem

Research questions and friction points this paper is trying to address.

Optimizing federated learning under device resource constraints
Adapting training parameters via Lagrangian dual optimization
Reducing memory and communication costs for edge devices
Innovation

Methods, ideas, or system contributions that make the work stand out.

Lagrangian dual optimization adapts hyperparameters dynamically
Freezing depth and gradient accumulation preserve training stability
Reduces memory usage and communication while maintaining performance
💼 Related Jobs
No related jobs found.
Dongqi Zheng
Dongqi Zheng
Apple; Purdue University, West Lafayette
W
Wenjin Fu
Carnegie Mellon University