Score
Designs and applies linear-algebra methods that decompose changes in model parameters or activations into a low-dimensional ‘‘essential’’ principal subspace and orthogonal residuals, estimate that subspace’s dimensionality (k_t) and principal directions, and quantify how much task-relevant signal is preserved by projection. Builds procedures for orthogonalized fusion and static model merging that combine or recover models by preserving high-energy activation directions and measuring compact decodability and recovery performance.
This work addresses the performance degradation commonly observed in multi-task model merging due to interference among tasks. The authors propose an efficient, training-free merging method that first constructs an intrinsic feature subspace dominated by task-specific parameter updates and projects individual task models onto this low-rank subspace for fusion. To further enhance knowledge retention and suppress redundancy, a multi-level polarized scaling mechanism is introduced to amplify critical parameters while attenuating less informative ones. By integrating principal component analysis, low-rank decomposition, and parameter projection, the approach substantially mitigates task interference across diverse task sets and model scales, preserving essential functionalities and achieving state-of-the-art performance in multi-task model merging.
This work addresses the lack of theoretical characterization in existing kernel methods for machine learning regarding the residual structure and energy stability of multichannel signals in complex systems. The authors propose an analytical framework grounded in operator defect identities, introducing the novel concept of “telescopic energy residuals.” By integrating iterative products with a λₙ-relaxed Kaczmarz scheme, they establish admissibility conditions for residuals and derive prior energy bounds. For the first time, this framework incorporates operator defect theory into kernel methods and kernel principal component analysis (KPCA), rigorously proving explicit convergence of generalized algorithms, a residual energy decomposition theorem, and stability criteria under noise. The approach significantly extends infinite-dimensional Kaczmarz theory to broader applications in machine learning.
This work addresses interference in model merging for multi-task learning, which often arises from conflicting parameter updates across tasks. The authors observe that output shifts induced by task-specific updates concentrate their energy within a low-dimensional “essential subspace” spanned by a few dominant directions. Building on this insight, they propose two training-free merging strategies: a static Essential Subspace Merging (ESM) method that mitigates interference through orthogonal fusion, and a dynamic variant (ESM++) that combines low-rank expert decomposition with prototype-based routing to enable task-adaptive integration. Experiments demonstrate that both approaches substantially alleviate inter-task interference across diverse task sets and model scales while effectively preserving individual task performance.
This paper addresses the weak theoretical foundations of matrix decomposition in machine learning by systematically constructing a self-consistent, comprehensive, and modern-application-oriented pedagogical framework. Methodologically, it grounds the exposition in numerical linear algebra and matrix analysis, unifying classical decompositions—including LU, QR, SVD, and block triangular factorizations—while integrating numerical stability analysis and Hermitian/Hilbert space theory. Crucially, it bridges traditional numerical analysis with deep learning’s backpropagation setting, emphasizing differentiability and computational robustness of decompositions in algorithm design and model optimization. The primary contribution is a compact, dual-purpose (teaching and research) knowledge system that fills critical gaps in both theoretical coherence and machine-learning relevance present in existing literature, thereby providing rigorous mathematical foundations for high-dimensional data modeling and efficient training.
In parametric dynamical systems, the Proper Orthogonal Decomposition (POD) basis drifts with parameters, degrading the accuracy of reduced-order models (ROMs). Method: This paper proposes the Projected Gaussian Process (pGP) framework—the first to formulate subspace adaptation as a statistical learning task mapping parameter space to the Grassmann manifold. It employs a two-stage geometric mapping: Euclidean space → horizontal space → Grassmann manifold, integrating POD, exponential/logarithmic maps, horizontal-space projection, and Gaussian process regression to enable uncertainty-aware POD subspace prediction while preserving manifold structure. Contribution/Results: Numerical experiments demonstrate that pGP significantly improves ROM accuracy and robustness in both parametric extrapolation and interpolation scenarios, and provides interpretable, calibrated confidence quantification—establishing a new paradigm for parameter-sensitive model reduction.
本文研究了SVD压缩对自注意力矩阵秩塌陷的影响,发现其在初始化时抑制但在预训练模型中加速秩塌陷,并解释了这一现象的原因。
This work addresses the limitation of conventional element-wise reconstruction error in tensor low-rank approximation, which fails to capture geometric degradation of multidimensional structures. Building upon the orthogonal Tucker model, the paper introduces a novel orthogonal decomposition of reconstruction error into directional loss—quantifying subspace deviations caused by truncation and noise—and interaction loss—measuring distortions in the multilinear interactions of the core tensor. A Wedin-type stability bound is established for the directional loss. Experiments on synthetic and hyperspectral data, leveraging matrix SVD, tensor Tucker decomposition, and subspace perturbation analysis, demonstrate that under identical reconstruction errors, directional loss can vary by up to 4.6×, and high directional loss strongly correlates with visual blurring, thereby underscoring the necessity of structure-aware error metrics.
Existing data-driven Koopman operator methods struggle to ensure approximate invariance of subspaces under the operator in non-Euclidean settings, limiting predictive accuracy. This work addresses this challenge by extending principal vector–guided subspace pruning to reproducing kernel Hilbert spaces (RKHS) for the first time. By precisely computing principal angles and vectors in RKHS, we introduce Kernel-SPV and its computationally efficient Nyström approximation–based variant, Approximate Kernel-SPV. These approaches overcome the limitations of traditional Euclidean formulations, significantly enhancing the invariance of Koopman-invariant subspaces while maintaining scalability and substantially improving prediction accuracy.
Standard boosting methods often suffer from redundancy among base learners due to repeatedly fitting correlated errors. This work proposes SCBoost, a novel framework that reformulates boosting from a geometric perspective. SCBoost introduces Spectral Residual Projection (SRP) to constrain each new learner to the orthogonal complement of the subspace spanned by previous predictions, ensuring it captures only previously unexplained information. Additionally, Covariance-Regularized Weighting (CRW) is employed to optimize ensemble weights, explicitly reducing inter-learner correlation. The approach enables an exact additive decomposition of residual energy and provably enhances the signal-to-noise ratio under isotropic noise assumptions. Empirical evaluations across ten benchmark datasets demonstrate that SCBoost significantly outperforms baseline methods, achieving particularly notable gains in accuracy and F1 score.
CORAM通过在奇异值分解框架下合并模型权重的正交旋转来解决模型合并问题,使用放大系数和分散切片技术以提高性能。