SOAEsV2-7B/72B: Full-Pipeline Optimization for State-Owned Enterprise LLMs via Continual Pre-Training, Domain-Progressive SFT and Distillation-Enhanced Speculative Decoding

📅 2025-05-07
📈 Citations: 0
Influential: 0
📄 PDF

career value

192K/year
🤖 AI Summary
Domestic large language models (LLMs) face three key challenges in state-owned asset management: limited model capacity, overreliance on supervised fine-tuning (SFT) data, and inefficient long-context reasoning. Method: We propose the SOAEsV2-7B/72B model series and an end-to-end optimization framework comprising: (1) a novel domain-specific progressive SFT curriculum learning paradigm that incrementally integrates general linguistic competence with domain expertise; (2) a distillation-enhanced speculative decoding architecture enabling accelerated long-context inference without compromising generation quality; and (3) a three-stage pipeline—continual pretraining, progressive fine-tuning, and efficient decoding—that jointly optimizes general capability retention and domain performance. Results: Experiments demonstrate 1.08× and 1.17× improvements in domain-task Rouge-1 and BLEU-4 scores, respectively; 1.39–1.52× inference speedup; and 99.8% preservation of general capabilities.

Technology Category

Application Category

📝 Abstract
This study addresses key challenges in developing domain-specific large language models (LLMs) for Chinese state-owned assets and enterprises (SOAEs), where current approaches face three limitations: 1) constrained model capacity that limits knowledge integration and cross-task adaptability; 2) excessive reliance on domain-specific supervised fine-tuning (SFT) data, which neglects the broader applicability of general language patterns; and 3) inefficient inference acceleration for large models processing long contexts. In this work, we propose SOAEsV2-7B/72B, a specialized LLM series developed via a three-phase framework: 1) continual pre-training integrates domain knowledge while retaining base capabilities; 2) domain-progressive SFT employs curriculum-based learning strategy, transitioning from weakly relevant conversational data to expert-annotated SOAEs datasets to optimize domain-specific tasks; 3) distillation-enhanced speculative decoding accelerates inference via logit distillation between 72B target and 7B draft models, achieving 1.39-1.52$ imes$ speedup without quality loss. Experimental results demonstrate that our domain-specific pre-training phase maintains 99.8% of original general language capabilities while significantly improving domain performance, resulting in a 1.08$ imes$ improvement in Rouge-1 score and a 1.17$ imes$ enhancement in BLEU-4 score. Ablation studies further show that domain-progressive SFT outperforms single-stage training, achieving 1.02$ imes$ improvement in Rouge-1 and 1.06$ imes$ in BLEU-4. Our work introduces a comprehensive, full-pipeline approach for optimizing SOAEs LLMs, bridging the gap between general language capabilities and domain-specific expertise.
Problem

Research questions and friction points this paper is trying to address.

Enhancing model capacity for knowledge integration and adaptability
Reducing reliance on domain-specific SFT data for broader applicability
Improving inference efficiency for large models in long contexts
Innovation

Methods, ideas, or system contributions that make the work stand out.

Continual pre-training integrates domain knowledge
Domain-progressive SFT employs curriculum-based learning
Distillation-enhanced speculative decoding accelerates inference
🔎 Similar Papers
No similar papers found.
J
Jingyang Deng
School of Mathematical Sciences and LMAM, Peking University, Beijing 100871, China
R
Ran Chen
School of Mathematical Sciences and LMAM, Peking University, Beijing 100871, China
Jo-Ku Cheng
Jo-Ku Cheng
Peking University
Artificial IntelligenceMachine Learning
J
Jinwen Ma
School of Mathematical Sciences and LMAM, Peking University, Beijing 100871, China