A Riemannian Geometry for Low-rank Adaptation

📅 2026-10-06
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the optimization inefficiency in Low-Rank Adaptation (LoRA) caused by parameter redundancy, which hinders weight updates from approximating full fine-tuning. To overcome this limitation, the authors construct a Riemannian metric tailored for LoRA based on quotient manifold theory and employ preconditioned gradient descent to eliminate redundant directions. Theoretically, they prove that this metric minimizes the Frobenius norm discrepancy between LoRA gradients and those of full fine-tuning, thereby enabling low-rank updates to precisely approximate full fine-tuning. Empirically, the proposed approach significantly enhances both the efficiency and performance of LoRA fine-tuning across language and vision tasks. Ultimately, this work establishes a rigorous Riemannian geometric theoretical foundation and introduces a novel paradigm for parameter-efficient fine-tuning.
📝 Abstract
Low-rank adaptation (LoRA) is widely used as a parameter-efficient fine-tuning technique for pre-trained deep neural networks, which approximates the weight update via full fine-tuning by a low-rank matrix $BA^\top$. This parameterization leads to the equivalence relation $(B, A) \sim (BG^{-1}, AG^\top)$ for any invertible matrix $G$ because $BA^\top = BG^{-1}(AG^\top)^\top$ and thus both pairs yield the same loss value. This relation induces a quotient manifold where matrices $(BG^{-1}, AG^\top)$ for all $G$ are identified, eliminating redundant directions along which the loss value remains unchanged. To respect the geometry of this manifold, the original search space is endowed with a Riemannian metric that is invariant under the equivalence relation. Such a metric induces preconditioning at each gradient step and ensures that each weight update via LoRA changes the loss value, leading to efficient optimization. In this paper, we propose a new Riemannian metric that is specifically tailored to LoRA to close the gap to full fine-tuning at the weight level. We theoretically show that LoRA with our preconditioning induced by this metric satisfies the following two properties at each iteration: (i) The weight update follows the direction closest to the gradient of full fine-tuning within the subspace of first-order weight changes allowed by the LoRA parameterization. (ii) The updated weight matrix is closer in Frobenius norm to that of full fine-tuning than the updated weight matrices of LoRA with conventional preconditioning and without preconditioning. These theoretical insights suggest that our preconditioning makes LoRA better approximate full fine-tuning, thereby leading to more efficient optimization. Experiments show the effectiveness and efficiency of our preconditioning for LoRA on fine-tuning tasks with language and vision domains.
Problem

Research questions and friction points this paper is trying to address.

Low-rank Adaptation
Riemannian Geometry
Parameter-efficient Fine-tuning
Quotient Manifold
Preconditioning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Low-rank Adaptation
Riemannian Geometry
Preconditioning
Quotient Manifold
Full Fine-tuning