Computational Limits of Low-Rank Adaptation (LoRA) for Transformer-Based Models

📅 2024-06-05
🏛️ arXiv.org
📈 Citations: 20
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the high computational complexity of gradient computation in Low-Rank Adaptation (LoRA) fine-tuning. Under the Strong Exponential Time Hypothesis (SETH), we establish the first computational phase-transition theory for LoRA gradient computation: we prove that a subquadratic approximation algorithm exists when the LoRA rank satisfies a norm-sharing upper-bound threshold; further, we design a chain-based low-rank approximation scheme enabling near-linear gradient updates. Methodologically, we propose a term-wise controlled analytical framework integrating fine-grained complexity analysis, hierarchical low-rank gradient approximation, and selective adaptation of attention weight subsets (Q/V/K). Our core contributions are: (1) the first characterization of the efficiency phase-transition threshold for LoRA gradient computation; (2) the identification of sufficient conditions for near-linear solvability; and (3) the first theory-driven acceleration pathway for low-rank adaptive methods.

Technology Category

Machine Learning: Mixture of Experts (MoE)Search and Optimization: Learning to SearchNatural Language Processing: Learning & Optimization for NLP

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
We study the computational limits of Low-Rank Adaptation (LoRA) update for finetuning transformer-based models using fine-grained complexity theory. Our key observation is that the existence of low-rank decompositions within the gradient computation of LoRA adaptation leads to possible algorithmic speedup. This allows us to (i) identify a phase transition behavior and (ii) prove the existence of nearly linear algorithms by controlling the LoRA update computation term by term, assuming the Strong Exponential Time Hypothesis (SETH). For the former, we identify a sharp transition in the efficiency of all possible rank-$r$ LoRA update algorithms for transformers, based on specific norms resulting from the multiplications of the input sequence $mathbf{X}$, pretrained weights $mathbf{W^star}$, and adapter matrices $alpha mathbf{B} mathbf{A} / r$. Specifically, we derive a shared upper bound threshold for such norms and show that efficient (sub-quadratic) approximation algorithms of LoRA exist only below this threshold. For the latter, we prove the existence of nearly linear approximation algorithms for LoRA adaptation by utilizing the hierarchical low-rank structures of LoRA gradients and approximating the gradients with a series of chained low-rank approximations. To showcase our theory, we consider two practical scenarios: partial (e.g., only $mathbf{W}_V$ and $mathbf{W}_Q$) and full adaptations (e.g., $mathbf{W}_Q$, $mathbf{W}_V$, and $mathbf{W}_K$) of weights in attention heads.
Problem

Research questions and friction points this paper is trying to address.

Analyze computational limits of LoRA for transformer fine-tuning
Identify phase transition in LoRA efficiency under SETH
Prove existence of linear algorithms via low-rank gradients
Innovation

Methods, ideas, or system contributions that make the work stand out.

Low-Rank Adaptation (LoRA) for transformer fine-tuning
Phase transition in LoRA efficiency via SETH
Almost linear algorithms using hierarchical low-rank structures
🔎 Similar Papers
Northwestern University | USTC | National Center for Theoretical Sciences | Adobe Research
Jerry Yao-Chieh Hu
Jerry Yao-Chieh Hu
Northwestern University
Machine Learning(* denotes equal contribution)
M
Maojiang Su
Department of Information and Computing Science, USTC, Hefei, Anhui 230026, China
E
En-Jui Kuo
Physics Division, National Center for Theoretical Sciences, Taipei 106319, Taiwan
Z
Zhao Song
Adobe Research, Seattle, WA 98103, USA
H
Han Liu
Department of Statistics and Data Science, Northwestern University, Evanston, IL 60208, USA