Score
Design and build pipelines that distill models or knowledge representations by using execution or runtime feedback to detect, prune, and correct erroneous or redundant entries and iteratively refine the distilled artifact. Integrate large language models as semantic arbiters to reconcile conflicting entries and consolidate redundant knowledge into a consistent, compact knowledge base or distilled model.
Enterprise-level Text-to-SQL deployment faces a trilemma among cost, security, and performance: large language models (LLMs) incur prohibitive inference costs, while small language models (SLMs) suffer from insufficient accuracy and robustness. This paper proposes Struct-SQL—the first structured chain-of-thought knowledge distillation framework grounded in query execution plans (QEPs). Rather than relying on unstructured natural-language reasoning, Struct-SQL uses QEPs as formal, executable reasoning blueprints to guide SLM training. The method integrates structured reasoning modeling, abstract QEP representation learning, and targeted SLM fine-tuning, thereby enhancing logical consistency and syntactic correctness of generated SQL. On standard benchmarks, Struct-SQL achieves an absolute +8.1% improvement in execution accuracy over unstructured CoT distillation baselines and significantly reduces syntax error rates. These results empirically validate the critical role of structured reasoning signals in enabling reliable semantic parsing for Text-to-SQL.
This work addresses the challenge that small-scale open-source language models underperform large closed-source models (e.g., GPT-3.5) on customer-service summarization tasks. To bridge this gap, we propose Analyze-Revise-Finetune (ARF), a targeted error-correcting distillation framework. ARF first systematically analyzes error patterns of a teacher model (Llama 3.1 70B), then employs a lightweight editing model to generate high-quality, task-specific correction data, and finally performs fine-grained fine-tuning on a student model (Llama 3.1 8B). ARF achieves, for the first time, statistically significant outperformance of an 8B-class open-source model over GPT-3.5 on customer-service summarization—while simultaneously reducing training costs and enabling sensitive-data localization. Its core innovation lies in tightly integrating error-driven data construction with knowledge distillation, establishing a scalable, privacy-preserving, and cost-efficient paradigm for lightweight model optimization.
研究通过跨模型错误提炼知识规则,提出Robust Failure, Conservative Repair方法,提高模型推理准确性。
This work investigates whether knowledge distillation (KD) is necessary and effective for bundle generation tasks powered by large language models (LLMs), aiming to sustain high performance at low computational cost. To this end, it systematically disentangles the impact of KD along three dimensions—format, knowledge volume, and utilization strategy—for the first time. The authors propose a hierarchical knowledge extraction framework, capturing pattern-level, rule-based, and deep-reasoning knowledge, coupled with a multi-paradigm adaptation mechanism supporting in-context learning (ICL), supervised fine-tuning (SFT), and hybrid fine-tuning. Through progressive knowledge extraction and multi-granularity quantization fusion, the approach significantly improves the efficiency–performance trade-off of student models: achieving ≥92% of teacher-model performance across multiple benchmarks while reducing inference cost by 67%. The core contribution lies in establishing the critical role of structured KD in bundle generation and delivering a reusable, lightweight deployment paradigm.
This study addresses the problems of world-knowledge inconsistency and systematic knowledge gaps in large language models (LLMs). To this end, we propose KonTest, a knowledge-graph-based automated consistency testing framework. Methodologically, we introduce a novel consistency verification paradigm integrating semantic-equivalent querying, metamorphic testing, and ontological oracles; additionally, we design a weighted LLM ensemble mechanism to mitigate knowledge omissions. Experimental evaluation across four mainstream LLMs demonstrates that KonTest identifies inconsistent inputs in 19.2% of test cases and detects an average of 16.5% knowledge gaps. Our ensemble strategy reduces knowledge gaps by 32.48%. Overall, KonTest provides a scalable, interpretable, and principled framework for assessing the reliability of LLMs’ factual knowledge—advancing both the rigor and transparency of LLM evaluation.
为解决基于大型语言模型的代理在处理特定项目问题时缺乏专业知识的问题,提出SkillForge框架,通过合成项目特有问题主动获取并提炼项目知识,提高问题解决效率。
This work addresses the inefficiency and unreliability of directly deploying large language models (LLMs) or their distilled variants for enterprise tasks, which are typically deterministic, structured, and heavily reliant on domain-specific knowledge under strict constraints of cost, latency, and reliability. To overcome these limitations, the authors propose a modular AI architecture that confines LLMs to structured information extraction while offloading knowledge storage and reasoning to dedicated knowledge bases and symbolic systems. This design circumvents the bottlenecks of monolithic models in terms of interpretability, reliability, and maintainability. Theoretical analysis and system implementation demonstrate that the proposed architecture offers a more efficient, transparent, and sustainable alternative to end-to-end LLM approaches, providing a scalable AI solution tailored for enterprise applications.
To address the suboptimal performance of small language models (SLMs) on resource-constrained edge devices, this paper proposes a systematic post-training framework integrating curriculum-based supervised fine-tuning (Curriculum SFT) and offline on-policy policy knowledge distillation. Notably, it is the first work to incorporate curriculum learning into the policy distillation pipeline, significantly enhancing the reasoning capability and task generalization of billion-parameter SLMs under stringent hardware constraints. Leveraging Ascend-based edge platform optimizations, the resulting instruction-tuned model achieves state-of-the-art (SOTA) performance across diverse complex tasks—matching the accuracy of large language models while reducing inference latency by over 60% and cutting hardware resource consumption by an order of magnitude. This work establishes a new, cost-effective, and generalizable paradigm for deploying high-performance SLMs in edge AI scenarios.
Knowledge distillation (KD) accelerates large-model training but poses risks of intellectual property leakage and model homogenization; existing detection methods—relying on identity or output similarity—are vulnerable to prompt engineering evasion. Method: We propose the first KD detection framework leveraging expert routing patterns in Mixture-of-Experts (MoE) architectures to construct structural-habit fingerprints, enabling efficient identification in both white-box and black-box settings. Innovatively, we model routing behavior as transferable “structural habits” and introduce Shadow-MoE—a technique that generalizes these habits to non-MoE models and black-box APIs. Contribution/Results: By integrating routing analysis, proxy MoE modeling, and multi-dimensional comparison, our method achieves >94% detection accuracy on standard benchmarks—significantly outperforming baselines—while demonstrating strong robustness against prompt-engineering attacks. This validates structural habits as traceable, general-purpose indicators of KD.
This study addresses the limitation of traditional knowledge distillation, which is confined to parameter imitation and fails to transfer the supra-parametric capabilities upon which agents rely, such as memory, tool use, and execution logic. We define agent distillation as the persistent transfer of task-solving knowledge and propose a taxonomic perspective based on knowledge retention loci, encompassing intra-model, artifact-based, execution framework, and cross-substrate dimensions. Furthermore, this work constructs a causal evaluation framework that disentangles transfer evidence from outcomes. By establishing a theoretical roadmap for multi-substrate knowledge transfer, this research lays the foundation for developing reliable, maintainable, and secure complex agent systems.