🤖 AI Summary
This study addresses the prevalent misconception that the rank parameter in LoRA functions as a capacity control mechanism, clarifying its true role under norm constraints. By integrating spectral analysis, distribution distance metrics, and Rademacher complexity theory, this work demonstrates that under a hard norm budget, the rank upper bound becomes inactive within the nuclear norm ball. Consequently, it redefines rank as a controlling factor for the reachable update set and spectral alignment cost. The primary contributions include establishing a rank-independent generalization complexity upper bound and deriving tight bounds on the minimum rank required for source-target task alignment. These findings provide a rigorous theoretical foundation for LoRA hyperparameter selection.
📝 Abstract
The rank $r$ in LoRA is widely treated as a capacity control: a smaller rank is assumed to yield a simpler model that generalizes better. We show that, under hard per-factor norm budgets---the idealization of the weight decay and norm control used in practice---this intuition breaks down. The reason is structural: under such budgets, the updates LoRA can reach are exactly the matrices of rank at most $r$ inside a nuclear-norm ball, and every complexity and displacement functional we analyze is maximized over this set by a rank-one update---so the rank cap never binds. The consequences follow directly. The linear-readout model class we study is identical for every $r \ge 1$, its Rademacher complexity carries no dependence on $r$, and the distance the adaptation can move the source distribution obeys a rank-independent upper bound that we show is sharp. If rank does not control capacity, where does it act? We identify two places. Statistically, replacing the per-factor budgets with a joint budget on the product restores a data-dependent, rank-sensitive complexity bound---though the gain appears only for well-spread feature distributions, and the worst case remains rank-free. Spectrally, rank sets the price of adaptation: canceling the leading singular directions of the pretrained weight requires both sufficient rank and sufficient budget. We bound the smallest rank achieving a desired source--target alignment, with upper and lower bounds that match under two-sided spectral decay. Together, these results recast rank as governing which updates are reachable and what cancellation costs---not how much capacity the model has.