train low-rank experts

Designs, trains, and evaluates model components that implement low-rank additive parameter updates (experts) attached to a shared backbone so the experts capture specialized features while preserving shared representations. This includes specifying low-rank parameterizations, training procedures and regularization to limit expert parameter growth and prevent interference with backbone features.

trainlow-rankexperts

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.03
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Low-Rank Adaptation for Foundation Models: A Comprehensive Review

Dec 31, 2024
MY
Menglin Yang
🏛️ Yale University | Nanyang Technological University | The Chinese University of Hong Kong | University of Electronic Science and Technology of China | Hong Kong University | Birla Institute of Technology and Science | Logs AI

To address the high computational cost and poor generalization in efficient adaptation of foundation models, this paper presents the first systematic survey of Low-Rank Adaptation (LoRA) extensions across broad classes of foundation models—including multimodal and scientific computing models. We propose a unified taxonomy that integrates matrix low-rank decomposition, modular adapter design, gradient-constrained optimization, and cross-task transfer analysis—thereby identifying key theoretical gaps and charting a new direction toward robustness-aware modeling. Covering over 100 state-of-the-art works, we uncover common mechanisms underlying LoRA’s cross-modal transferability and pinpoint critical deployment bottlenecks. Our synthesis delivers a methodological framework and reproducible implementation pathways for lightweight adaptation of general-purpose foundation models, advancing efficient, robust, and scalable model customization paradigms.

Computational Resource ReductionEfficient Fine-tuningLarge-scale Pre-trained Models

Must-Read Papers

Most classic and influential ideas
View more

This work addresses a critical vulnerability in parameter-level defenses against unauthorized expert integration during model merging. Existing defenses fail because the small magnitude of task vectors allows pre-trained weights to dominate the merged model, inadvertently acting as static anchors that compromise security. The study is the first to expose this anchoring risk and introduces Anchor-Guided Attack (AGA), which leverages task vector analysis and linear transformation modeling to reconstruct and invert the transformation matrix, thereby bypassing prevailing defenses. To counter this threat, the authors propose Anchor-Repulsive Fine-tuning (ARF), a novel defense strategy that actively repels such anchoring effects. Experimental results demonstrate that AGA reliably circumvents both single and composite defenses, while ARF effectively mitigates the attack, restoring robustness to model merging pipelines.

free-ridingmodel mergingparameter-level defenses

From Parameter to Representation: A Closed-Form Approach for Controllable Model Merging

Nov 14, 2025
JW
Jialin Wu
🏛️ Rocket Force University of Engineering | Xidian University

In multi-task model merging, parameter interference impedes users from flexibly trading off performance across tasks according to personal preferences. Existing “compile-then-query” approaches rely on costly offline multi-objective optimization, whose computational complexity grows exponentially with the number of tasks. Method: We propose a representation correction paradigm that bypasses parameter-space optimization entirely and instead directly rectifies the final-layer representations of merged models. We design a user-preference-aware optimal linear transformation, enabling architecture-agnostic, single-step, closed-form solution. Contribution/Results: Our method reduces computational complexity from exponential to linear in the number of tasks. Experiments demonstrate that it enables instantaneous generation of Pareto-optimal models, achieving superior Pareto frontier quality, more precise preference alignment, and significantly lower computational cost compared to prior methods.

Current methods have exponential complexity growth with task numbersExisting approaches require costly offline multi-objective optimization processesModel merging faces parameter interference in multitask performance optimization

Why Do More Experts Fail? A Theoretical Analysis of Model Merging

May 27, 2025
ZW
Zijing Wang
🏛️ Northeastern University | LMU Munich

This work addresses the scalability bottleneck in model merging—specifically, the performance degradation observed as the number of experts increases. We establish, for the first time, a theoretical framework grounded in Gaussian width and approximate kinematics, revealing parameter-space saturation as the fundamental limiting factor; we further prove that performance gains exhibit strictly concave decay and admit a unique optimal merging threshold. Building on this insight, we propose Reparameterized Heavy-Tailed (RHT) merging, which alleviates saturation constraints via heavy-tailed reparameterization of expert weights. Extensive evaluation across 12 knowledge-intensive and general-purpose benchmarks demonstrates that RHT significantly delays performance decay and raises the upper bound for multi-task fusion. The implementation is open-sourced. To our knowledge, this is the first theoretically grounded paradigm for scalable model merging, offering provable guarantees on convergence behavior and capacity limits.

Introduces method to enhance merged model performanceInvestigates scalability limits of merging multiple expert modelsProves upper bound on model merging due to parameter constraints

Prompt learning fails in polluted Mixture-of-Experts (MoE) models due to two key mechanisms: parameter coupling between the pretrained backbone and prompt experts—causing prompt weights to vanish—and algebraic interactions governed by partial differential equations that induce learning slowdown. Method: We introduce a distinguishability condition to decouple parameter dynamics, systematically characterize how expert architecture—e.g., sparsity and overlap—affects estimation convergence rates, and derive matching minimax lower bounds. Results: We establish tight convergence rates theoretically, providing the first quantitative explanation of prompt learning failure within the minimax framework. Numerical experiments empirically validate both prompt vanishing and convergence slowdown. Our analysis yields principled theoretical foundations and practical guidance for designing robust prompt-based MoE systems.

Address prompt vanishing and parameter interaction issues.Analyze convergence in contaminated mixture of experts.Investigate expert structures and their convergence effects.

Maintaining Structural Integrity in Parameter Spaces for Parameter Efficient Fine-tuning

May 23, 2024
CS
Chongjie Si
🏛️ Shanghai Jiao Tong University | Shanghai AI Laboratory

To address structural distortion and topological inconsistency in high-dimensional parameter spaces (e.g., 4D tensors) induced by low-rank approximation in parameter-efficient fine-tuning, this paper proposes a structure-preserving low-rank core space modeling method. Unlike conventional low-rank adapters (e.g., LoRA), which are restricted to linear weight matrices, our approach explicitly models and preserves the intrinsic topological structure of the original high-dimensional parameter space—achieving compact and accurate reconstruction of N-dimensional parameter updates via high-order tensor decomposition. Evaluated across CV, NLP, and multimodal benchmarks, the method yields an average accuracy improvement of 1.8% under identical parameter budgets, while reducing structural distortion by 37%, significantly outperforming existing baselines.

Enable parameter-efficient fine-tuning across diverse dimensional spacesModel changes via low-rank core space with consistent topologyPreserve structural integrity in high-dimensional parameter spaces

Latest Papers

What's happening recently
View more

Existing dynamic model merging approaches suffer from suboptimal parameter allocation between shared and expert modules, struggling to balance accuracy and efficiency. This work proposes DiDi-Merging, a novel framework that introduces differentiable rank allocation into dynamic merging for the first time, enabling efficient and compact multi-task models by optimizing the parameter budget of low-rank modules. The method integrates data-free distillation to recover task fidelity and supports dynamic expert activation. Remarkably, DiDi-Merging matches the performance of current methods using only 1.24× the parameters of a single fine-tuned model and surpasses them at 1.4×, substantially reducing the storage overhead compared to other approaches that typically require more than 2× the base model size.

accuracy-efficiency trade-offdynamic model merginglow-rank adaptation

This work addresses the high deployment cost of Mixture-of-Experts (MoE) models caused by their massive expert parameters, a challenge inadequately resolved by existing compression methods that struggle to balance accuracy and scalability. The authors propose an efficient compression approach that preserves the original router and leverages functional co-activation patterns among experts to cluster them. Within each cluster, one full-precision dominant expert is retained, while others are represented as low-rank corrections. Furthermore, they introduce BTExperts, a tree-structured organization enabling computation sharing during inference. Evaluated on Qwen3-30B-A3B and Gemma-4-26B-A4B, the method achieves approximately 50% expert compression while outperforming baseline models in downstream accuracy and perplexity across most tasks, with performance gains increasing as the number of experts scales.

expert compressionlow-rank decompositionMixture-of-Experts

This work addresses the performance degradation of task-specific experts in multitask model merging caused by parameter interference, as well as the high inference cost and storage overhead of existing dynamic methods that rely on redundant expert copies. The authors propose ReTeX, a framework that models parameter interference as an affine transformation of expert parameters and approximates it with a learnable additive offset, enabling a single merged model to recover near-original expert performance. Innovatively, ReTeX introduces a router-free task identifier that leverages singular value decomposition (SVD) subspace projection residuals to match task identities, achieving the first subspace-based task recognition without additional storage. Experiments demonstrate that ReTeX recovers over 95% of standalone expert performance across vision and NLP tasks and exhibits strong generalization and adaptive knowledge interpolation capabilities on unseen tasks.

inference efficiencymulti-task model mergingout-of-distribution generalization

This study addresses the storage and deployment bottlenecks caused by parameter redundancy in large Mixture-of-Experts (MoE) models, as well as the irreducible errors in existing pruning-merging methods arising from routing and expert heterogeneity. We propose SLBF, a data-free weight reconstruction framework that establishes the first structural error bounds for pruning and merging. By introducing shared low-rank factorization and post-hoc canonical fixation, SLBF achieves efficient cross-expert compression without requiring original training data while fully preserving routing mechanisms. Evaluated across five MoE architectures ranging from 16B to 122B parameters, SLBF demonstrates lower reconstruction error and faster convergence, comprehensively outperforming three mainstream compression approaches.

Data-freeLarge Language ModelsMixture-of-Experts

This study addresses the limitation of relying solely on expert utilization rates to assess removal damage during expert pruning in Mixture-of-Experts (MoE) models, emphasizing the necessity of preserving functional substitutability to maintain output distributions. To this end, it proposes a training-free expert pruning framework that introduces a consensus residual-based scoring mechanism for evaluating functional substitutability. By integrating an exact single-deletion identity with calibration token aggregation, the method achieves precise pruning without requiring gradients or recovery training. Extensive experiments across multiple large-scale models and varying pruning ratios demonstrate that the proposed approach attains the highest macro-average score over nine evaluation tasks. It significantly outperforms the REAP baseline while effectively reducing reverse KL divergence, highlighting its efficacy in maintaining model performance under aggressive pruning conditions.

Expert PruningFunctional ReplaceabilityLarge Language Models

Hot Scholars

RS

Robert Sim

Sr. Principal Research Manager, Microsoft
machine learningresponsible AIprivacyrobotics
CB

Chetan Bansal

Microsoft
AI AgentsDistributed SystemsSoftware Engineering
JD

Jiawen Deng

University of Electronic Science and Technology of China
NLPAI SafetyAffective Computing
FR

Fuji Ren

Professor of University of Electronic Science and Technology of China
Artificial IntelligenceComputer ScienceAffective Computing
XZ

Xuchao Zhang

Principal Researcher @ Microsoft
NLPData MiningMachine Learning