Score
Designs, trains, and evaluates model components that implement low-rank additive parameter updates (experts) attached to a shared backbone so the experts capture specialized features while preserving shared representations. This includes specifying low-rank parameterizations, training procedures and regularization to limit expert parameter growth and prevent interference with backbone features.
To address the high computational cost and energy consumption in large language model training and fine-tuning, this paper systematically uncovers the dual role of low-rank structure throughout optimization: (i) low-rankness spontaneously emerges in gradient dynamics, and (ii) its implicit regularization governs the generalization properties of the converged solution. We establish the first unified theoretical framework linking the dynamical origin of low-rankness to convergence behavior, thereby bridging the theoretical foundations of LoRA and masked training. Our approach integrates gradient low-rank decomposition, optimization dynamics modeling, and implicit regularization analysis to provide rigorous theoretical justification for parameter-efficient fine-tuning. Experiments demonstrate that our method significantly reduces computational overhead and energy consumption while preserving model performance. This work advances low-rank adaptation from an empirical practice to a principle-driven paradigm.
To address the high computational cost and poor generalization in efficient adaptation of foundation models, this paper presents the first systematic survey of Low-Rank Adaptation (LoRA) extensions across broad classes of foundation models—including multimodal and scientific computing models. We propose a unified taxonomy that integrates matrix low-rank decomposition, modular adapter design, gradient-constrained optimization, and cross-task transfer analysis—thereby identifying key theoretical gaps and charting a new direction toward robustness-aware modeling. Covering over 100 state-of-the-art works, we uncover common mechanisms underlying LoRA’s cross-modal transferability and pinpoint critical deployment bottlenecks. Our synthesis delivers a methodological framework and reproducible implementation pathways for lightweight adaptation of general-purpose foundation models, advancing efficient, robust, and scalable model customization paradigms.
This work addresses a critical vulnerability in parameter-level defenses against unauthorized expert integration during model merging. Existing defenses fail because the small magnitude of task vectors allows pre-trained weights to dominate the merged model, inadvertently acting as static anchors that compromise security. The study is the first to expose this anchoring risk and introduces Anchor-Guided Attack (AGA), which leverages task vector analysis and linear transformation modeling to reconstruct and invert the transformation matrix, thereby bypassing prevailing defenses. To counter this threat, the authors propose Anchor-Repulsive Fine-tuning (ARF), a novel defense strategy that actively repels such anchoring effects. Experimental results demonstrate that AGA reliably circumvents both single and composite defenses, while ARF effectively mitigates the attack, restoring robustness to model merging pipelines.
In multi-task model merging, parameter interference impedes users from flexibly trading off performance across tasks according to personal preferences. Existing “compile-then-query” approaches rely on costly offline multi-objective optimization, whose computational complexity grows exponentially with the number of tasks. Method: We propose a representation correction paradigm that bypasses parameter-space optimization entirely and instead directly rectifies the final-layer representations of merged models. We design a user-preference-aware optimal linear transformation, enabling architecture-agnostic, single-step, closed-form solution. Contribution/Results: Our method reduces computational complexity from exponential to linear in the number of tasks. Experiments demonstrate that it enables instantaneous generation of Pareto-optimal models, achieving superior Pareto frontier quality, more precise preference alignment, and significantly lower computational cost compared to prior methods.
This work addresses the scalability bottleneck in model merging—specifically, the performance degradation observed as the number of experts increases. We establish, for the first time, a theoretical framework grounded in Gaussian width and approximate kinematics, revealing parameter-space saturation as the fundamental limiting factor; we further prove that performance gains exhibit strictly concave decay and admit a unique optimal merging threshold. Building on this insight, we propose Reparameterized Heavy-Tailed (RHT) merging, which alleviates saturation constraints via heavy-tailed reparameterization of expert weights. Extensive evaluation across 12 knowledge-intensive and general-purpose benchmarks demonstrates that RHT significantly delays performance decay and raises the upper bound for multi-task fusion. The implementation is open-sourced. To our knowledge, this is the first theoretically grounded paradigm for scalable model merging, offering provable guarantees on convergence behavior and capacity limits.
Prompt learning fails in polluted Mixture-of-Experts (MoE) models due to two key mechanisms: parameter coupling between the pretrained backbone and prompt experts—causing prompt weights to vanish—and algebraic interactions governed by partial differential equations that induce learning slowdown. Method: We introduce a distinguishability condition to decouple parameter dynamics, systematically characterize how expert architecture—e.g., sparsity and overlap—affects estimation convergence rates, and derive matching minimax lower bounds. Results: We establish tight convergence rates theoretically, providing the first quantitative explanation of prompt learning failure within the minimax framework. Numerical experiments empirically validate both prompt vanishing and convergence slowdown. Our analysis yields principled theoretical foundations and practical guidance for designing robust prompt-based MoE systems.
To address structural distortion and topological inconsistency in high-dimensional parameter spaces (e.g., 4D tensors) induced by low-rank approximation in parameter-efficient fine-tuning, this paper proposes a structure-preserving low-rank core space modeling method. Unlike conventional low-rank adapters (e.g., LoRA), which are restricted to linear weight matrices, our approach explicitly models and preserves the intrinsic topological structure of the original high-dimensional parameter space—achieving compact and accurate reconstruction of N-dimensional parameter updates via high-order tensor decomposition. Evaluated across CV, NLP, and multimodal benchmarks, the method yields an average accuracy improvement of 1.8% under identical parameter budgets, while reducing structural distortion by 37%, significantly outperforming existing baselines.
Existing dynamic model merging approaches suffer from suboptimal parameter allocation between shared and expert modules, struggling to balance accuracy and efficiency. This work proposes DiDi-Merging, a novel framework that introduces differentiable rank allocation into dynamic merging for the first time, enabling efficient and compact multi-task models by optimizing the parameter budget of low-rank modules. The method integrates data-free distillation to recover task fidelity and supports dynamic expert activation. Remarkably, DiDi-Merging matches the performance of current methods using only 1.24× the parameters of a single fine-tuned model and surpasses them at 1.4×, substantially reducing the storage overhead compared to other approaches that typically require more than 2× the base model size.
This work addresses the high deployment cost of Mixture-of-Experts (MoE) models caused by their massive expert parameters, a challenge inadequately resolved by existing compression methods that struggle to balance accuracy and scalability. The authors propose an efficient compression approach that preserves the original router and leverages functional co-activation patterns among experts to cluster them. Within each cluster, one full-precision dominant expert is retained, while others are represented as low-rank corrections. Furthermore, they introduce BTExperts, a tree-structured organization enabling computation sharing during inference. Evaluated on Qwen3-30B-A3B and Gemma-4-26B-A4B, the method achieves approximately 50% expert compression while outperforming baseline models in downstream accuracy and perplexity across most tasks, with performance gains increasing as the number of experts scales.
This work addresses the performance degradation of task-specific experts in multitask model merging caused by parameter interference, as well as the high inference cost and storage overhead of existing dynamic methods that rely on redundant expert copies. The authors propose ReTeX, a framework that models parameter interference as an affine transformation of expert parameters and approximates it with a learnable additive offset, enabling a single merged model to recover near-original expert performance. Innovatively, ReTeX introduces a router-free task identifier that leverages singular value decomposition (SVD) subspace projection residuals to match task identities, achieving the first subspace-based task recognition without additional storage. Experiments demonstrate that ReTeX recovers over 95% of standalone expert performance across vision and NLP tasks and exhibits strong generalization and adaptive knowledge interpolation capabilities on unseen tasks.
This study addresses the storage and deployment bottlenecks caused by parameter redundancy in large Mixture-of-Experts (MoE) models, as well as the irreducible errors in existing pruning-merging methods arising from routing and expert heterogeneity. We propose SLBF, a data-free weight reconstruction framework that establishes the first structural error bounds for pruning and merging. By introducing shared low-rank factorization and post-hoc canonical fixation, SLBF achieves efficient cross-expert compression without requiring original training data while fully preserving routing mechanisms. Evaluated across five MoE architectures ranging from 16B to 122B parameters, SLBF demonstrates lower reconstruction error and faster convergence, comprehensively outperforming three mainstream compression approaches.
This study addresses the limitation of relying solely on expert utilization rates to assess removal damage during expert pruning in Mixture-of-Experts (MoE) models, emphasizing the necessity of preserving functional substitutability to maintain output distributions. To this end, it proposes a training-free expert pruning framework that introduces a consensus residual-based scoring mechanism for evaluating functional substitutability. By integrating an exact single-deletion identity with calibration token aggregation, the method achieves precise pruning without requiring gradients or recovery training. Extensive experiments across multiple large-scale models and varying pruning ratios demonstrate that the proposed approach attains the highest macro-average score over nine evaluation tasks. It significantly outperforms the REAP baseline while effectively reducing reverse KL divergence, highlighting its efficacy in maintaining model performance under aggressive pruning conditions.