π€ AI Summary
This study addresses the challenge of joint nonlinear feature learning with asymmetric kernel matrices across multiple data sources by proposing the eKSVD method. This approach extends KSVD to multi-source scenarios, integrating explicit neural network mappings with a covariance framework to overcome the limitations of traditional Mercer kernels. Furthermore, it generalizes Lanczosβ theorem through Lagrangian dual optimization based on KKT conditions, enabling the joint learning of multi-source projections and pairwise coupling. Experimental results demonstrate that eKSVD significantly outperforms baseline methods in multi-source data processing, exhibiting high flexibility and effectiveness.
π Abstract
Kernel Singular Value Decomposition (KSVD) learns a pair of singular vectors w.r.t. an asymmetric kernel matrix, which can be induced by two data sources, e.g., the queries and keys in self-attention or the rows and columns of a given matrix. In this work, we extend KSVD to multiple data sources, namely eKSVD, which conducts joint nonlinear feature learning upon asymmetric kernels. In the primal formulation, the projections associated with each data source are jointly learned to capture maximal information, while incorporating pair-wise couplings. With the Lagrangian and its Karush-Kuhn-Tucker (KKT) conditions, the optimization in the dual leads to a generalization of the shifted eigenvalue problem in Lanczos decomposition theorem of KSVD. Further, a covariance-based framework is derived together with using neural networks (NNs) for explicit feature mappings, complementary to the kernel-based interpretation and optimization. Numerical experiments verify the effectiveness of our eKSVD compared to methods based on Mercer kernels for tackling multiple data sources, and our innovation of deploying NNs demonstrates great flexibility for kernel methods.