Targeted Error Correction in Knowledge Distillation: Small Language Models Surpass GPT

📅 2025-11-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge that small-scale open-source language models underperform large closed-source models (e.g., GPT-3.5) on customer-service summarization tasks. To bridge this gap, we propose Analyze-Revise-Finetune (ARF), a targeted error-correcting distillation framework. ARF first systematically analyzes error patterns of a teacher model (Llama 3.1 70B), then employs a lightweight editing model to generate high-quality, task-specific correction data, and finally performs fine-grained fine-tuning on a student model (Llama 3.1 8B). ARF achieves, for the first time, statistically significant outperformance of an 8B-class open-source model over GPT-3.5 on customer-service summarization—while simultaneously reducing training costs and enabling sensitive-data localization. Its core innovation lies in tightly integrating error-driven data construction with knowledge distillation, establishing a scalable, privacy-preserving, and cost-efficient paradigm for lightweight model optimization.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: SummarizationSearch and Optimization: Learning to Search

Application Category

Search and Retrieval-Augmented AI: Large language models for searchUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
📝 Abstract
We introduce an Analyze-Revise-Finetune (ARF) pipeline that enables smaller open-source language models (LLMs) to surpass substantially larger proprietary models in customer service summarization tasks. The pipeline first analyzes and categorizes common errors in summaries produced by a teacher model (GPT-3.5), then performs a targeted revision using a compact editor model (Llama 3.1 70B) to generate high-quality, refined training data. Fine-tuning a smaller student model (Llama 3.1 8B) on this refined data resulted in superior summarization performance compared to GPT-3.5. The ARF pipeline improves cost efficiency and data privacy while maintaining competitive accuracy, illustrating a generalizable framework for enhancing open-source LLMs across diverse downstream applications.
Problem

Research questions and friction points this paper is trying to address.

Improving small language models' summarization accuracy
Reducing dependency on large proprietary AI models
Enhancing data privacy and cost efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Analyze-Revise-Finetune pipeline corrects teacher model errors
Compact editor model generates refined training data
Fine-tuning smaller student model surpasses larger models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Hee-Jin Lee
eBay Inc. 2025 Hamilton Avenue, San Jose, CA, USA
Z
Zhen Guo
eBay Inc. 2025 Hamilton Avenue, San Jose, CA, USA
L
Luchao Jin
eBay Inc. 2025 Hamilton Avenue, San Jose, CA, USA
M
M. Goudarzi
eBay Inc. 2025 Hamilton Avenue, San Jose, CA, USA