low-rank subspace alignment

Designs and analyzes algorithms, regularizers, and evaluation metrics that align or increase overlap between low-rank parameter subspaces (singular/vector subspaces) of model weight updates or adapter factors by measuring singular-subspace similarity and encouraging local updates toward a reference subspace. Builds subspace-constrained or subspace-regularized federated LoRA training and aggregation procedures that constrain or regularize LoRA factors across clients to stabilize aggregated updates and reduce client-level divergence.

low-ranksubspacealignment

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.01
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Heterogeneous client data in federated LoRA leads to misalignment of low-rank update subspaces, causing aggregation conflicts and slow convergence. This work is the first to reveal the critical impact of such subspace misalignment on federated LoRA performance and proposes a subspace regularization method that explicitly aligns clients’ geometric structures by constraining local updates to remain close to a shared global reference subspace. Leveraging SVD-based basis alignment analysis, the proposed approach significantly outperforms baseline methods such as FedAvg on RoBERTa-large, achieving an accuracy of 0.429 ± 0.011 and a subspace basis overlap of 0.9999, thereby demonstrating the effectiveness of high-precision subspace alignment.

federated learninggeometric alignmentlow-rank adaptation

This study addresses the instability of shared structures in federated LoRA, where parameter similarity is susceptible to initialization interference and neglects local input distributions. To overcome this, we propose FedSAIL, a framework that exploits second-order moment statistics of input-aware action matrices to mine robust shared subspace geometry across clients. Our analysis reveals that the principal singular directions of action matrices are more robust than direct parameter alignment. Accordingly, we establish a novel aggregation paradigm that replaces conventional weight averaging with subspace-guided aggregation, complemented by regularization constraints to optimize low-rank adaptation training. Extensive experiments demonstrate that FedSAIL significantly improves predictive performance across multiple benchmarks while substantially reducing communication overhead.

Client HeterogeneityFederated LearningInput-aware

In federated learning with heterogeneous clients fine-tuning large models via LoRA, data-parameter interference leads to unstable adapter aggregation. This work is the first to formulate this issue as a subspace alignment problem and introduces Dysco, a dynamic subspace augmentation method. Dysco extracts data-agnostic subspace bases from local activations each round, enabling the server to construct client-specific merged subspaces optimized for compatibility and mitigate representation drift through a multi-round augmentation mechanism. Theoretical analysis yields a closed-form solution for a fixed server-side merged subspace. Experiments demonstrate that Dysco reduces training loss by up to 9× over baselines on synthetic tasks and improves the performance of five federated algorithms by up to 4.3% on MIMIC-IV clinical text classification, with only a 0.9% increase in runtime overhead, significantly outperforming existing federated LoRA approaches.

Client HeterogeneityFederated LearningLow-Rank Adaptation

This work addresses the limitations of independent low-rank projections in federated learning, which cause subspace fragmentation, aggregation bias, and substantial memory and communication overhead. We propose a shared gradient low-rank projection framework that achieves exact aggregation in low-rank coordinates through intra-round basis sharing, inter-round dynamic refreshing, and a projected SGD algorithm. This approach theoretically eliminates projection bias and overcomes fixed rank-budget constraints. Experiments demonstrate that the proposed method outperforms LoRA baselines and matches full-parameter fine-tuning on vision-language tasks while significantly reducing communication costs and optimizer state memory.

Communication EfficiencyFederated LearningFoundation Models

Selective Aggregation for Low-Rank Adaptation in Federated Learning

Oct 02, 2024
PG
Pengxin Guo
🏛️ The University of Hong Kong | Stanford University | Shenyang Institute of Automation | Chinese Academy of Sciences

In federated learning (FL), jointly optimizing shared general knowledge and client-specific knowledge remains challenging. Method: This paper first uncovers the semantic specialization of the A/B matrices in LoRA and its variants (rsLoRA, VeRA): the A matrix encodes low-rank, cross-client shared general knowledge, while the B matrix captures client-specific deviations. Building on this insight, we propose FedSA-LoRA—a novel FL framework that selectively aggregates only the A matrices, enabling explicit decoupling and targeted fusion of knowledge. Contribution/Results: FedSA-LoRA establishes the first general-purpose, LoRA-family-oriented federated low-rank adaptation paradigm. Extensive experiments on NLU and NLG tasks demonstrate its superiority over baselines in accuracy, faster convergence, stronger generalization, and reduced communication overhead. The implementation is publicly available.

Extends FedSA-LoRA to rsLoRA and VeRA variantsInvestigates LoRA in federated learning via matrix asymmetry analysisProposes FedSA-LoRA to share only A matrices for aggregation

Latest Papers

What's happening recently
View more

This work addresses the trade-off between performance and efficiency in parameter-efficient fine-tuning of large language models by systematically reinterpreting Low-Rank Adaptation (LoRA) through the lens of signal processing. Leveraging classical low-rank modeling and inverse problem theory, it establishes a unified framework to understand both existing and future efficient fine-tuning methods. The study proposes a three-dimensional technical framework encompassing architecture design, optimization strategies, and full-lifecycle deployment, integrating core techniques such as singular value decomposition, rank expansion, cross-layer tensorization, norm-invariant optimization, and parameterization-aware solvers. This approach provides theoretical grounding and principled design guidelines for LoRA and its variants, while extending their applicability across pre-training, post-training, and deployment stages, thereby fostering bidirectional integration between signal processing and deep learning.

Architectural DesignFoundation ModelsLow-Rank Adaptation

Existing LoRA merging methods often suffer from parameter interference in shared parameter spaces, leading to degraded generation performance. This work proposes Subspace Signal Routing (SSR), which reframes LoRA fusion as a signal routing problem. SSR constructs a unified subspace by concatenating along the rank dimension, decouples mixed signals using an inverse correlation matrix, and allocates purified signals to task-specific subspaces via a direction-guided matrix. Theoretically, this approach is shown to be equivalent to the ordinary least squares solution. Leveraging the additivity of sufficient statistics, SSR further introduces a streaming algorithm that enables efficient online updates. Experimental results demonstrate that SSR significantly outperforms current LoRA merging techniques while maintaining computational efficiency, achieving notable improvements in multi-task generation quality.

diffusion modelsLoRA mergingparameter interference

This work addresses the gradient collapse issue in high-rank LoRA during federated fine-tuning of large language models, which arises from amplified statistical variance due to multi-client aggregation. Existing scaling strategies overlook the interaction between federated aggregation and adapter rank. To resolve this, we propose SFed-LoRA, a novel framework that theoretically characterizes the impact of federated aggregation on LoRA rank for the first time and derives an optimal scaling factor dependent on both the number of clients and the adapter rank. This scaling effectively corrects aggregation-induced errors and stabilizes training without modifying model architecture or increasing inference overhead. Extensive experiments demonstrate that SFed-LoRA significantly improves convergence speed and stability of high-rank LoRA across diverse tasks, models, and heterogeneous data settings, outperforming current state-of-the-art baselines.

Client HeterogeneityFederated LearningGradient Collapse

Existing methods struggle to effectively merge multiple task-specific LoRA adapters into a single low-rank adapter without causing capability fragmentation or violating the low-rank structure. This work proposes a novel “Compress-then-Merge” (CtM) paradigm: it first constructs a shared r-dimensional subspace from the LoRA weights and orthogonally projects each adapter onto this subspace, then performs standard merging within the resulting r×r core coordinate space, followed by truncated SVD to strictly enforce the target rank constraint. Evaluated across multiple models and tasks, CtM significantly outperforms existing single-LoRA baselines and substantially narrows the performance gap with full-parameter merging, achieving for the first time an efficient and rank-preserving LoRA fusion.

adapter compressionLoRAlow-rank adaptation

This work addresses the challenges of knowledge interference and parameter conflicts arising from sequential unlearning in continual machine unlearning. To mitigate these issues, the authors propose a static orthogonal constraint mechanism that leverages SVD-guided orthogonal subspace projection to strictly confine each LoRA update to the orthogonal complement of previously learned task subspaces. This approach achieves task isolation without requiring dynamic routing. Empirical results demonstrate that the method effectively preserves retained knowledge during repeated unlearning iterations, achieving a retention accuracy of 58.1% after 30 rounds of continual unlearning on CIFAR-100 and MNIST—substantially outperforming existing static fusion baselines (12.7%) while maintaining efficient unlearning capability.

continual machine unlearningknowledge preservationparameter collision

Hot Scholars

JF

Junfeng Fang

National University of Singapore
Model EditingAI SafetyLLM ExplainabilityAI4Science
TL

Trung Le

Faculty of Information Technology, Monash University, Australia
Adversarial Machine LearningGenerative ModelsModel UnlearningModel Editing
MH

Mehrtash Harandi

Department of Electrical and Computer Systems Engineering, Monash University
Machine LearningComputer Vision
DP

Dinh Phung

Professor, Monash University
machine learningdeep learningoptimal transportprobabilistic inference
GH

Guangyan Huang

Deakin University
stream data miningintelligent sensingsensor datatext analytics