🤖 AI Summary
This work addresses a critical limitation in low-rank knowledge distillation: output-level distillation fails to explicitly align the low-rank subspaces of teacher and student models, leading to subspace misalignment and reduced compression efficiency. To resolve this, the paper introduces a spectral alignment mechanism that jointly optimizes three sources of error—subspace misalignment, coefficient mismatch, and irreducible residual—through data-weighted student subspace reference updates and a differentiable principal angle loss. The proposed method integrates LoRA adaptation, subspace projection, and data-weighted spectral decomposition. Empirical results demonstrate that it reduces subspace misalignment error from 51% to nearly zero on synthetic tasks. On six GLUE benchmarks, it outperforms the strongest spectral baseline on five tasks at rank r=4 and achieves state-of-the-art performance on SST-2 and CoLA at r=8.
📝 Abstract
Distilling a fine-tuned teacher into a LoRA-adapted student is a standard recipe for parameter-efficient compression, but output-level KD does not explicitly control which rank-$r$ weight subspace the adapter occupies. We propose \textbf{SAD-LoRA} (\textbf{S}pectral \textbf{A}lignment \textbf{D}istillation), which selects this subspace from the data-weighted student-space reference update $\DWT\Sigx^{1/2}$ and maintains it during training via a differentiable principal-angle loss on $\colspan(B)$. We show that the data-weighted distillation error decomposes exactly into subspace misalignment, within-subspace coefficient mismatch, and irreducible rank residual; standard KD can affect the first term only indirectly through output gradients. On controlled synthetic problems with a flat teacher spectrum, SAD-LoRA reduces the subspace-misalignment term from $51\%$ to nearly zero and lifts final subspace alignment from $0.49$ to $1.00$. On RoBERTa-large to RoBERTa-base distillation across six GLUE tasks, SAD-LoRA improves rank efficiency: at $r{=}4$, it matches or beats the strongest included spectral baseline on five of six tasks, and at $r{=}8$ it gives the best result on SST-2 and CoLA. Ablations identify subspace alignment as the load-bearing component, while coefficient matching is auxiliary.