PIVOT: Perplexity-Informed KD-to-RL Transition Scheduling for Vertical-Domain Few-Shot Distillation

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of fixed distillation-to-reinforcement-learning transition strategies in few-shot classification for vertical domains, which overlook sample-level heterogeneity. To this end, we propose PIVOT, a dynamic framework featuring a novel per-sample routing mechanism based on teacher model perplexity. Unlike conventional global fixed switching strategies, PIVOT adaptively allocates individual samples to either online policy distillation (OPD) or GRPO-based reinforcement learning, thereby achieving a dynamic balance between knowledge acquisition and reward refinement. Experimental results demonstrate that PIVOT significantly outperforms existing baselines on benchmarks such as Banking77, effectively enhancing downstream performance while improving training stability.
📝 Abstract
Vertical-domain few-shot classification remains challenging for small language models, as limited supervision makes it difficult to acquire domain-specific decision knowledge. On-Policy Distillation (OPD) can improve teacher-guided adaptation by supervising student-generated rollouts, while GRPO-based reinforcement learning can further refine downstream predictions. However, existing KD-to-RL pipelines typically rely on globally fixed transition schedules, ignoring that different samples may require different amounts of teacher-guided acquisition before reward-driven refinement. We propose PIVOT (Perplexity-Informed Transition Optimization), a dynamic transition framework that routes samples between OPD and GRPO according to teacher-evaluated sequence perplexity. PIVOT moves low-perplexity samples to GRPO for reward-driven refinement while keeping high-perplexity samples under OPD for continued domain knowledge acquisition. Experiments on Banking77 and HWU64 show that PIVOT consistently outperforms continued OPD and globally synchronized OPD$\rightarrow$GRPO baselines under the same number of post-warm-up student optimization steps, achieving stronger downstream performance and more stable training dynamics.
Problem

Research questions and friction points this paper is trying to address.

few-shot classification
vertical-domain
knowledge distillation
reinforcement learning
transition scheduling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Few-Shot Distillation
Perplexity-Informed Routing
On-Policy Distillation
GRPO
Dynamic Transition Scheduling
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Heng Li
Heng Li
University of Science and Technology of China
Robotics
Y
Yong Zhang
Ping An Technology (Shenzhen) Co., Ltd., China
Ning Cheng
Ning Cheng
TeraHop
Z
Zhigen Li
Ping An Technology (Shenzhen) Co., Ltd., China
Y
Yun Zhu
Ping An Technology (Shenzhen) Co., Ltd., China
Y
Yanmeng Wang
Ping An Technology (Shenzhen) Co., Ltd., China
Shaojun Wang
Shaojun Wang
Soochow University, TU/e, University of Strasbourg
NanophotonicsLight-matter interactionsNanofabrication
J
Jing Xiao
Ping An Technology (Shenzhen) Co., Ltd., China