🤖 AI Summary
This study addresses the challenges of knowledge obsolescence and difficulty in injecting proprietary information into large language models (LLMs). We systematically investigate how task typology affects knowledge retention during injection. Using diverse architectures—including Llama, Qwen, and Phi—we conduct comparative fine-tuning across understanding-oriented tasks (e.g., question answering, cloze) and mapping-oriented tasks (e.g., translation, text-to-JSON). We introduce a dual-axis evaluation framework measuring both knowledge retention rate and cross-context transferability. Our key findings are: (1) depth of cognitive engagement—not data volume—is the primary determinant of injection efficacy; (2) understanding-oriented tasks achieve 48% retention, significantly outperforming mapping-oriented tasks (17–20%), a trend consistent across architectures and aligned with scaling laws; and (3) all models suffer >35% performance degradation in novel contexts, indicating that injected knowledge remains superficial and fails to integrate semantically.
📝 Abstract
As the knowledge of large language models (LLMs) becomes outdated over time, there is a growing need for efficient methods to update them, especially when injecting proprietary information. Our study reveals that comprehension-intensive fine-tuning tasks (e.g., question answering and blanks) achieve substantially higher knowledge retention rates (48%) compared to mapping-oriented tasks like translation (17%) or text-to-JSON conversion (20%), despite exposure to identical factual content. We demonstrate that this pattern persists across model architectures and follows scaling laws, with larger models showing improved retention across all task types. However, all models exhibit significant performance drops when applying injected knowledge in broader contexts, suggesting limited semantic integration. These findings show the importance of task selection in updating LLM knowledge, showing that effective knowledge injection relies not just on data exposure but on the depth of cognitive engagement during fine-tuning.