Score
Design, derive, and implement Jacobian matrices and their determinants (including analytic derivations and efficient numerical/log-determinant estimators) for differentiable multivariate mappings, with support for per-layer or input–output mappings in parameterized models. Analyze those Jacobians to evaluate local invertibility and volume change, test stability via determinant sign, detect shared local linear operations, and quantify functional similarity or interpretable internal transformations between units or components.
This paper addresses the analytical computation of first- and second-order derivatives—specifically, the Jacobian and Hessian matrices—of the Kullback–Leibler (KL) divergence between multivariate Gaussian distributions with respect to their mean vectors and symmetric positive-definite covariance matrices. Leveraging matrix differential calculus, we unify the Magnus differential framework with Minka’s technique to derive, for the first time, closed-form expressions for all first- and second-order derivatives of the KL divergence with respect to both parameters. The derivation is mathematically rigorous, fully transparent, and pedagogically interpretable. The resulting formulae substantially enhance the computability and interpretability of gradient and curvature information in variational inference, information geometry, and probabilistic model optimization. By providing exact, compact second-order derivatives, this work furnishes a critical mathematical foundation for developing efficient, curvature-aware optimization algorithms in statistical machine learning.
Efficient and accurate computation of higher-order derivatives of generalized eigenvalues/vectors for symmetric matrices and generalized singular values/vectors for rectangular matrices—parameterized linearly or nonlinearly—is essential yet theoretically underdeveloped, especially under degeneracy and practical constraints. Method: We rigorously derive closed-form expressions for first- and second-order analytical derivatives using matrix differential calculus, validate Jacobian and Hessian matrices via numerical differentiation, and implement an open-source R package supporting diverse degenerate cases and application-specific constraints. Contribution/Results: This work establishes the first unified, complete high-order differential theory framework for these spectral decompositions. The proposed method substantially improves both computational efficiency and numerical accuracy of derivative evaluation. It has been successfully applied to multivariate statistical modeling and covariance structure optimization, providing a rigorous theoretical foundation and practical computational infrastructure for parameter sensitivity analysis and gradient-based optimization in eigen-decomposition–dependent models.
Current deep learning frameworks lack autograd-compatible fractional-order matrix differentiation, hindering the integration of fractional calculus into neural network optimization. To address this, we propose the fractional-order Jacobian matrix differential ( J^alpha ) and establish a matrix-form fractional chain rule compatible with automatic differentiation, enabling end-to-end differentiable propagation of fractional-order gradients through hidden layers. Leveraging PyTorch, we design a Riemann–Liouville-based fractional linear layer (FLinear) that seamlessly integrates into standard backpropagation. Experiments on multilayer perceptrons demonstrate that FLinear significantly reduces training loss and improves test accuracy, while maintaining reasonable GPU memory consumption and training efficiency. This work provides both a theoretical foundation and a practical implementation for incorporating fractional calculus into differentiable programming paradigms for deep learning.
To address the failure of gradient descent in multi-objective optimization caused by objective conflicts, this paper proposes the Jacobian Descent (JD) algorithm, which iteratively updates parameters directly in the Jacobian matrix space of the vector-valued objective function. Key contributions include: (1) a novel gradient projection mechanism that rigorously eliminates inter-objective conflicts while preserving the relative influence of each objective proportional to its gradient norm; (2) the first formalization of the instance-wise risk minimization (IWRM) paradigm, treating the loss of each training sample as an independent optimization objective; and (3) an efficient Gramian matrix approximation to substantially reduce memory and computational overhead in Jacobian computation. Theoretical analysis establishes JD’s improved convergence guarantees over standard methods. Empirical evaluation on image classification tasks demonstrates that IWRM consistently outperforms conventional average-loss minimization, validating both the efficacy and scalability of JD.
This paper addresses the weak theoretical foundations of matrix decomposition in machine learning by systematically constructing a self-consistent, comprehensive, and modern-application-oriented pedagogical framework. Methodologically, it grounds the exposition in numerical linear algebra and matrix analysis, unifying classical decompositions—including LU, QR, SVD, and block triangular factorizations—while integrating numerical stability analysis and Hermitian/Hilbert space theory. Crucially, it bridges traditional numerical analysis with deep learning’s backpropagation setting, emphasizing differentiability and computational robustness of decompositions in algorithm design and model optimization. The primary contribution is a compact, dual-purpose (teaching and research) knowledge system that fills critical gaps in both theoretical coherence and machine-learning relevance present in existing literature, thereby providing rigorous mathematical foundations for high-dimensional data modeling and efficient training.
This work addresses the geometric constraints and computational complexity inherent in generative modeling of symmetric positive-definite (SPD) matrices by introducing “reverse telescoping mapping”—a novel unconstrained coordinate system that decouples SPD matrices into interpretable volume (log-determinant) and shape (relative scale and off-diagonal covariance) components. The proposed representation yields a lossless symbolic encoding of both a matrix and its inverse, reducing the complexity of key operations to O(p²) in the transformed domain and enabling linear interpolation along unit-determinant geodesics. Leveraging this framework, we develop volume–shape disentangled conditional flow matching and intrinsic diffusion models, successfully synthesizing bimodal SPD distributions, generating fMRI brain connectomes, and performing manifold-intrinsic diffusion tasks at scale (p = 200).
This work addresses the long-standing issue of inconsistency in generalized matrix inverses under nonsingular diagonal transformations—a problem particularly critical in robotics, tracking, and control systems where physical unit sensitivity is paramount. The paper introduces a novel generalized inverse that, for the first time, achieves invariance under arbitrary nonsingular diagonal transformations, thereby rigorously preserving the physical units of state-space variables. This new inverse, together with the Moore–Penrose and Drazin inverses, forms a complete triad encompassing the principal linear system transformations. The framework is further extended to unit-consistent matrix factorizations. Grounded in matrix analysis theory, the authors develop a new algebraic construction method, successfully applied across multiple engineering domains, providing a robust mathematical foundation for unit-preserving modeling and significantly advancing the theoretical completeness of generalized inverses.
This work addresses the gap between abstract Riemannian geometry and practical algorithmic implementation by systematically developing a computationally tractable geometric framework for Riemannian optimization. Focusing on canonical matrix manifolds—Stiefel, Grassmann, and symmetric positive-definite (SPD) manifolds—it explicitly derives core geometric structures, including tangent spaces, metric tensors, Levi-Civita connections, curvature operators, and geodesics, all expressed in coordinate- and matrix-based forms amenable to numerical computation. Furthermore, it provides closed-form expressions for the Riemannian gradient, Hessian, exponential map, and retraction operators. To the best of our knowledge, this is the first unified formulation that translates classical differential-geometric constructions into a consistent, implementation-ready framework, thereby bridging theory and practice and offering a rigorous foundation for efficient and accurate algorithm design in Riemannian optimization and geometric machine learning.
This work investigates the identification of shared and view-specific latent structures in multi-view data, focusing on scenarios where symmetric positive definite matrices exhibit common eigendirections. By analyzing the spectral properties of matrix interpolants of the form \(A^{1-x}B^x\), the study establishes that the operator norm displays strict or approximate log-linear behavior if and only if the matrices share dominant eigenvectors. Leveraging this theoretical insight, the authors develop a multi-manifold learning framework that integrates complex matrix interpolation, spectral analysis, and singular vector alignment to effectively disentangle common and specific latent variables across views. This approach provides a novel theoretical foundation and a practical methodology for multi-view representation learning.
This work investigates the generalization ability of nonlinear least-squares models with ridge regularization at local minima attained after training. Leveraging average algorithmic stability analysis, it characterizes the local geometry of parameter space through the empirical Jacobian Gram matrix and a residual-curvature term. For the first time, an effective dimension is defined with respect to the trained model rather than the initialization point, and combined with curvature and covering complexity to derive a generalization error bound that depends on the geometric structure of learned features rather than the number of parameters. The framework explicitly links this bound to intrinsic data manifold dimensions or the number of stable activation regions in ReLU networks. Experiments confirm Jacobian contraction during training, the tightness of residual-curvature linearization, and a strong alignment between the proposed bound and observed generalization gaps.