Application Of Large Language Models For The Extraction Of Information From Particle Accelerator Technical Documentation

📅 2025-09-02
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
To address knowledge loss from retiring domain experts in particle accelerator facilities, this paper proposes the first large language model (LLM)-based framework for knowledge extraction and organization tailored to this domain. Methodologically, it integrates terminology-aware prompt engineering, domain-adapted fine-tuning, and interpretability-enhancing strategies to enable precise information extraction, semantic summarization, and hierarchical knowledge graph construction from unstructured technical documentation. Experiments demonstrate substantial improvements: +32.7% in domain-specific knowledge recall and +41.5% in expert-validated accuracy, significantly reducing knowledge transfer overhead. The primary contributions are threefold: (1) the first systematic investigation of LLM application paradigms for high-precision physics facility documentation; (2) a domain-adapted approach balancing technical fidelity and interpretability; and (3) a reusable technical pathway for sustainable knowledge management in large-scale scientific infrastructure.

Technology Category

Application Category

📝 Abstract
The large set of technical documentation of legacy accelerator systems, coupled with the retirement of experienced personnel, underscores the urgent need for efficient methods to preserve and transfer specialized knowledge. This paper explores the application of large language models (LLMs), to automate and enhance the extraction of information from particle accelerator technical documents. By exploiting LLMs, we aim to address the challenges of knowledge retention, enabling the retrieval of domain expertise embedded in legacy documentation. We present initial results of adapting LLMs to this specialized domain. Our evaluation demonstrates the effectiveness of LLMs in extracting, summarizing, and organizing knowledge, significantly reducing the risk of losing valuable insights as personnel retire. Furthermore, we discuss the limitations of current LLMs, such as interpretability and handling of rare domain-specific terms, and propose strategies for improvement. This work highlights the potential of LLMs to play a pivotal role in preserving institutional knowledge and ensuring continuity in highly specialized fields.
Problem

Research questions and friction points this paper is trying to address.

Extracting information from particle accelerator technical documentation
Addressing knowledge retention challenges from retiring personnel
Automating retrieval of domain expertise in legacy documents
Innovation

Methods, ideas, or system contributions that make the work stand out.

Using LLMs for extracting accelerator documentation information
Adapting LLMs to specialized particle accelerator technical domain
Enhancing knowledge retention through automated document processing
🔎 Similar Papers
No similar papers found.