Which and When to Admit: Gradient Admission for Data-Centric Small Language Model Finetuning

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses gradient conflicts, the failure of static data selection, and update overwriting caused by low-rank subspace saturation during LoRA fine-tuning of small language models. To overcome these challenges, we propose GRADE, a framework that decouples data selection from gradient admission for the first time. It introduces a state-aware selector coupled with a real-time aligned multi-task gradient field filtering mechanism. Furthermore, a self-calibrating gating module dynamically regulates the timing and direction of gradients entering the low-rank subspace, effectively preventing destructive overwriting and enhancing optimization trajectory coherence. Extensive experiments demonstrate that GRADE significantly outperforms baseline methods across multiple mainstream architectures, establishing itself as the only approach that consistently improves standard LoRA across all evaluated models.
📝 Abstract
LoRA fine-tuning adapts small language models (SLMs) to heterogeneous instruction data within a low-rank update subspace, making it vulnerable to three structural problems: conflicting gradients that cancel, static data selection that cannot track evolving learning dynamics, and subspace saturation that causes later updates to overwrite useful directions. We argue that effective adaptation therefore requires controlling which data-induced gradients enter the LoRA subspace and when. We propose GRADE (GRadient-Aligned Data-centric rEcipe), a data-centric framework combining two mechanisms: a state-aware selector that continually admits samples aligned with the evolving multi-task gradient field, and a self-calibrating step-level gate that rejects updates likely to cause destructive overwrite near saturation. Across three current-generation backbones and a heterogeneous seven-dataset instruction pool, GRADE outperforms strong data-selection and PEFT-stabilization baselines in accuracy and robustness. It is the only method to improve consistently over standard LoRA on every architecture, while producing more coherent gradient trajectories and less destructive overwrite. These results show that successful SLM adaptation depends not only on which data are selected, but also on which gradients are allowed to enter and persist in the constrained update subspace.
Problem

Research questions and friction points this paper is trying to address.

Small Language Models
LoRA Fine-tuning
Gradient Conflict
Data Selection
Subspace Saturation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Gradient Admission
LoRA Fine-tuning
State-aware Selector
Step-level Gate
Data-centric Framework