Score
Designs and implements processes and systems to add, inject, or refresh factual information in an existing model or knowledge store, including pipelines for curating or generating training data (e.g., paraphrases, question–answer pairs) and selecting update targets. Builds and evaluates update mechanisms — such as fine-tuning, parameter-efficient adapters, or data-injection methods — that integrate new knowledge while preserving the model's general capabilities and avoiding unwanted interference.
To address the problem of knowledge obsolescence and inefficient updating in large language models (LLMs), this paper proposes the first taxonomy of knowledge editing that jointly considers *mechanisms* (e.g., parameter modification, external memory) and *functions* (e.g., factual, temporal, conceptual knowledge), formally defining edit objectives and corresponding evaluation tasks. Unlike prior taxonomies focusing solely on editing mechanisms, our framework systematically incorporates knowledge functionality as a core analytical dimension. Through a comprehensive survey of 120+ editing methods, we construct a unified framework spanning editing strategies, knowledge categories, and evaluation benchmarks. Our analysis reveals systematic trade-offs: distinct knowledge functions critically influence editing robustness, generalization, and side-effect profiles. The work clarifies the applicability boundaries of existing approaches, identifies *functional misalignment*—the mismatch between editing mechanisms and target knowledge functionality—as a fundamental bottleneck, and highlights key open challenges, including interpretable editing, cross-functional transfer, and dynamic knowledge evolution.
This paper addresses three core challenges in large language models (LLMs): difficulty in dynamically expanding knowledge, weak integration of heterogeneous multi-source knowledge, and insufficient long-term consistency guarantees. To tackle these, we propose the first unified analytical framework that systematically integrates four complementary paradigms: continual learning, parametric model editing, retrieval-augmented generation (RAG), and implicit preference modeling. We introduce a structured taxonomy of knowledge types—factual, domain-specific, linguistic, and preference-based—and characterize their evolution along three dimensions: consistency, scalability, and verifiability, thereby constructing a comprehensive methodology map for knowledge expansion. Our key contributions include: (1) establishing a cross-paradigm evaluation benchmark; (2) advocating modular, composable adaptation design; and (3) advancing standardization in evaluation protocols. The framework provides both theoretical foundations and practical guidelines for developing evolvable, trustworthy, and scenario-adaptive knowledge-enhanced LLMs. (149 words)
This study addresses the challenges of knowledge obsolescence and difficulty in injecting proprietary information into large language models (LLMs). We systematically investigate how task typology affects knowledge retention during injection. Using diverse architectures—including Llama, Qwen, and Phi—we conduct comparative fine-tuning across understanding-oriented tasks (e.g., question answering, cloze) and mapping-oriented tasks (e.g., translation, text-to-JSON). We introduce a dual-axis evaluation framework measuring both knowledge retention rate and cross-context transferability. Our key findings are: (1) depth of cognitive engagement—not data volume—is the primary determinant of injection efficacy; (2) understanding-oriented tasks achieve 48% retention, significantly outperforming mapping-oriented tasks (17–20%), a trend consistent across architectures and aligned with scaling laws; and (3) all models suffer >35% performance degradation in novel contexts, indicating that injected knowledge remains superficial and fails to integrate semantically.
This work investigates the intrinsic relationship between task customization and knowledge injection in language model fine-tuning, identifying the root causes of knowledge injection difficulty. Using the Gemini v1.5 series, we construct a synthetically designed continuous-spectrum dataset and conduct controlled experiments with fine-grained evaluation. We quantitatively isolate three critical factors—question-answer format, information granularity, and multi-step reasoning capability—for the first time. Results show that QA-format training significantly enhances knowledge generalization; numerical knowledge is more prone to forgetting than categorical knowledge; fine-tuned models struggle with multi-step reasoning; and injection difficulty does not differ significantly between stylistic and factual knowledge. Crucially, the study challenges the conventional dichotomy between task customization and knowledge injection, demonstrating their mechanistic unity. This provides both theoretical grounding and practical guidance for controllable knowledge injection in LLMs.
To address the challenges of inefficient knowledge injection, severe catastrophic forgetting, and degradation of general capabilities in low-resource continual pretraining (CPT) of large language models (LLMs), this paper proposes a novel knowledge injection paradigm based solely on instruction tuning. Methodologically, it pioneers the use of small models to synthesize high-information-density, multi-hop reasoning–enhanced instruction data for targeted knowledge updates; further, it integrates knowledge distillation with retrieval-augmented modeling to jointly strengthen factual memory retention and preserve general reasoning and instruction-following abilities. A dedicated benchmark—Companies—is introduced to rigorously evaluate knowledge injection efficacy. Experimental results demonstrate that our approach significantly outperforms conventional CPT under extremely limited data budgets, achieving breakthroughs in mitigating forgetting, improving factual accuracy, enhancing complex contextual understanding, and enabling robust retrieval integration.
Existing in-context editing (ICE) methods struggle to disentangle newly injected knowledge from the model’s intrinsic parametric knowledge, leading to knowledge conflicts and inconsistent multi-hop reasoning. To address this, we propose a lightweight, parameter-free decoupled ICE framework that explicitly isolates knowledge injection from native inference via reasoning-path masking. Our approach integrates prompt engineering, dynamic path-mask generation, external knowledge retrieval, and in-model self-verification into an end-to-end interpretable editing pipeline. Crucially, it employs a hybrid retrieval–LLM self-verification mechanism to jointly ensure factual accuracy and reasoning-path consistency. Evaluated on multi-hop question answering benchmarks, our method significantly outperforms state-of-the-art ICE approaches, achieving simultaneous improvements in editing accuracy and reasoning consistency. The implementation is publicly available.
To address performance degradation and diminished general capabilities in large language models (LLMs) during sequential knowledge editing, this paper proposes a two-stage knowledge updating framework. First, robust supervised fine-tuning (R-SFT) internalizes new knowledge; second, the fine-tuned model is parameter-space fused with the original base model. This work introduces the novel “fine-tuning + fusion” paradigm—requiring no architectural modifications—enabling high-accuracy sequential editing while preserving pre-edit capabilities. Evaluated under a rigorous knowledge editing benchmark across multiple rounds of sequential edits, our method achieves a 23.5% improvement in knowledge correction accuracy and constrains performance decay on original tasks to within 0.8%, substantially outperforming state-of-the-art approaches.
In the era of large language models, traditional record-centric data engineering struggles to meet the demand for organizational knowledge as executable infrastructure. This work proposes a novel paradigm—knowledge architecture—that systematically reimagines core data engineering mechanisms by upgrading ETL, data lineage, and catalogs into knowledge ingestion, change detection, provenance, and knowledge catalogs. It introduces knowledge views and a three-tier layered model (raw–refined–operational) to structure knowledge effectively. By integrating emerging standards such as LLM Wiki and Open Knowledge Format (OKF), this study formally defines knowledge architecture for the first time and establishes a theoretical framework that supports knowledge representation, governance, and operational delivery, enabling direct invocation of organizational knowledge by humans, agents, workflows, and models alike.
This work addresses the lack of a unified evaluation framework for knowledge graph integration pipelines, which hinders systematic comparison and selection of methods. To bridge this gap, the paper introduces KGI-Bench, the first comprehensive benchmark specifically designed for evaluating knowledge graph data integration. KGI-Bench assesses integration performance across three key dimensions—coverage, correctness, and consistency—when incorporating heterogeneous input data (structured, semi-structured, and unstructured) into a target knowledge graph. Using a curated dataset in the movie domain, the benchmark evaluates twelve representative integration pipelines, revealing significant performance variations attributable to input data types and architectural choices. The results demonstrate the effectiveness and practical utility of KGI-Bench in enabling rigorous, reproducible evaluation of knowledge graph integration approaches.
This work addresses the limited generalization capability of existing knowledge editing methods for large language models and their difficulty in handling unstructured knowledge. The authors propose CoT2Edit, a novel paradigm that introduces instruction-guided chain-of-thought (CoT) reasoning into knowledge editing for the first time. CoT2Edit leverages language model agents to automatically generate CoT instruction data and integrates supervised fine-tuning (SFT), Group Relative Policy Optimization (GRPO), and retrieval-augmented generation (RAG) to enable unified editing and application of both structured and unstructured knowledge. Evaluated on three open-source large language models, CoT2Edit achieves significant improvements in generalization across six diverse knowledge editing scenarios after only a single round of training.
This work addresses the challenge of efficiently transferring domain-specific knowledge from a fine-tuned model to a new base architecture when the original proprietary training data is inaccessible. The authors propose an automated, data-free knowledge distillation method that identifies critical knowledge regions by analyzing discrepancies in perplexity between the fine-tuned and base models. Leveraging only a small set of representative prompts, the approach generates synthetic training data and combines iterative prompt expansion with parameter-efficient fine-tuning (e.g., LoRA) to effectively transfer knowledge. Notably, the method operates without any auxiliary discriminator and achieves state-of-the-art performance across multiple tasks, demonstrating high accuracy and flexibility in data-free knowledge transfer.
This work addresses the limitations of existing knowledge editing methods, which predominantly focus on atomic facts and struggle to support coherent multi-context reasoning or generalize effectively. The authors propose reframing knowledge internalization as a reasoning problem rather than a memorization task. Their approach introduces new knowledge through synthetically constructed background stories and automatically generates multi-hop questions to train models in multi-step reasoning. By integrating knowledge distillation, a student model learns to emulate the teacher’s reasoning process without direct access to the newly introduced knowledge. Built upon three core principles—background story construction, multi-step reasoning, and knowledge distillation—the method significantly enhances the model’s ability to integrate and flexibly apply new knowledge in complex reasoning scenarios.