🤖 AI Summary
This work addresses the challenge that small-scale open-source language models underperform large closed-source models (e.g., GPT-3.5) on customer-service summarization tasks. To bridge this gap, we propose Analyze-Revise-Finetune (ARF), a targeted error-correcting distillation framework. ARF first systematically analyzes error patterns of a teacher model (Llama 3.1 70B), then employs a lightweight editing model to generate high-quality, task-specific correction data, and finally performs fine-grained fine-tuning on a student model (Llama 3.1 8B). ARF achieves, for the first time, statistically significant outperformance of an 8B-class open-source model over GPT-3.5 on customer-service summarization—while simultaneously reducing training costs and enabling sensitive-data localization. Its core innovation lies in tightly integrating error-driven data construction with knowledge distillation, establishing a scalable, privacy-preserving, and cost-efficient paradigm for lightweight model optimization.
📝 Abstract
We introduce an Analyze-Revise-Finetune (ARF) pipeline that enables smaller open-source language models (LLMs) to surpass substantially larger proprietary models in customer service summarization tasks. The pipeline first analyzes and categorizes common errors in summaries produced by a teacher model (GPT-3.5), then performs a targeted revision using a compact editor model (Llama 3.1 70B) to generate high-quality, refined training data. Fine-tuning a smaller student model (Llama 3.1 8B) on this refined data resulted in superior summarization performance compared to GPT-3.5. The ARF pipeline improves cost efficiency and data privacy while maintaining competitive accuracy, illustrating a generalizable framework for enhancing open-source LLMs across diverse downstream applications.