SAD-LoRA: Spectral Alignment for Low-Rank Knowledge Distillation

📅 2026-07-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses a critical limitation in low-rank knowledge distillation: output-level distillation fails to explicitly align the low-rank subspaces of teacher and student models, leading to subspace misalignment and reduced compression efficiency. To resolve this, the paper introduces a spectral alignment mechanism that jointly optimizes three sources of error—subspace misalignment, coefficient mismatch, and irreducible residual—through data-weighted student subspace reference updates and a differentiable principal angle loss. The proposed method integrates LoRA adaptation, subspace projection, and data-weighted spectral decomposition. Empirical results demonstrate that it reduces subspace misalignment error from 51% to nearly zero on synthetic tasks. On six GLUE benchmarks, it outperforms the strongest spectral baseline on five tasks at rank r=4 and achieves state-of-the-art performance on SST-2 and CoLA at r=8.
📝 Abstract
Distilling a fine-tuned teacher into a LoRA-adapted student is a standard recipe for parameter-efficient compression, but output-level KD does not explicitly control which rank-$r$ weight subspace the adapter occupies. We propose \textbf{SAD-LoRA} (\textbf{S}pectral \textbf{A}lignment \textbf{D}istillation), which selects this subspace from the data-weighted student-space reference update $\DWT\Sigx^{1/2}$ and maintains it during training via a differentiable principal-angle loss on $\colspan(B)$. We show that the data-weighted distillation error decomposes exactly into subspace misalignment, within-subspace coefficient mismatch, and irreducible rank residual; standard KD can affect the first term only indirectly through output gradients. On controlled synthetic problems with a flat teacher spectrum, SAD-LoRA reduces the subspace-misalignment term from $51\%$ to nearly zero and lifts final subspace alignment from $0.49$ to $1.00$. On RoBERTa-large to RoBERTa-base distillation across six GLUE tasks, SAD-LoRA improves rank efficiency: at $r{=}4$, it matches or beats the strongest included spectral baseline on five of six tasks, and at $r{=}8$ it gives the best result on SST-2 and CoLA. Ablations identify subspace alignment as the load-bearing component, while coefficient matching is auxiliary.
Problem

Research questions and friction points this paper is trying to address.

knowledge distillation
LoRA
low-rank adaptation
subspace alignment
spectral alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Spectral Alignment
Low-Rank Adaptation
Knowledge Distillation
Principal-Angle Loss
LoRA
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
O
Omer Tariq
Neubility Inc, Seoul, South Korea
S
Syed Muhammad Raza
Neubility Inc, Seoul, South Korea
J
Jeongbae Son
Neubility Inc, Seoul, South Korea