Overcoming Prior Barriers: Supervised Fine-Tuning under Long-Tail Distribution

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of rare concepts forming "prior barriers" during large model fine-tuning due to insufficient pre-training support, which constrains long-tail knowledge acquisition. We propose PASS, a method that formally defines prior barriers and derives theoretical risk bounds to quantify learning difficulty under long-tailed distributions. Based on this formulation, we design an evidence-strength-based adaptive instruction selection algorithm that dynamically allocates fine-tuning budgets to overcome learning bottlenecks for high-barrier concepts. Experimental results demonstrate that PASS comprehensively outperforms seven state-of-the-art methods across four settings and significantly surpasses uniform allocation strategies. This work provides effective theoretical and methodological foundations for mitigating long-tail forgetting in large models.
📝 Abstract
Supervised fine-tuning (SFT) adapts pretrained large language models (LLMs) to downstream tasks, but the required concepts can receive substantially different levels of pretrained support. Frequent concepts are more likely to be well learned, whereas rare concepts may remain weakly represented. We introduce a novel notion named prior barrier to quantify how strongly the pretrained model supports competing concepts over the target concept. We observe that prior barriers follow a long-tail distribution, placing head and tail concepts at different starting points for SFT: head concepts face lower prior barriers, whereas tail concepts require additional instructions to overcome their higher prior barriers. Our theoretical analysis further derives a predictive risk bound for SFT under long-tail prior barriers, explicitly characterizing how the prior barrier and accumulated SFT evidence jointly determine predictive performance. Motivated by this prior barrier-dependent demand, we propose PASS, an adaptive SFT instruction selection method that constructs reference-derived concepts and estimates the distinguishing evidence provided by each instruction, and adaptively allocates the selection budget toward concepts that remain insufficiently covered under the current selection. In this way, PASS jointly considers which instructions can provide useful evidence and where additional supervision is needed under a limited budget. Experiments show that our method consistently outperforms seven state-of-the-art instruction selection methods on four backbone-budget settings. An ablation study further shows that PASS's adaptive allocation consistently improves over uniform allocation.
Problem

Research questions and friction points this paper is trying to address.

Supervised Fine-Tuning
Prior Barrier
Long-Tail Distribution
Instruction Selection
Large Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Prior Barrier
Supervised Fine-Tuning
Long-Tail Distribution
Instruction Selection
Adaptive Allocation
🔎 Similar Papers
No similar papers found.