A Cybersecurity MLPS Large Language Model with Multi-Path Retrieval Fusion

πŸ“… 2026-07-27
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitations of current manual and rule-based approaches to compliance analysis under China’s Multi-Level Protection Scheme (MLPS) for cybersecurity, which struggle to deliver stable and consistent intelligent judgments in complex scenarios. To overcome this, the authors propose a domain-specific large language model tailored for MLPS that innovatively integrates a multi-path retrieval mechanism combining hierarchical tree structures with token-level matching. This design ensures comprehensive rule coverage while suppressing irrelevant contextual interference. Furthermore, a multidimensional weighted scoring system is introduced to enable precise citation of regulatory clauses and traceable reasoning. Experimental results across ten representative tasks demonstrate that the proposed method significantly outperforms baseline models in clause-level accuracy, interpretability of conclusions, and practical deployability.
πŸ“ Abstract
The Multi-Level Protection Scheme (MLPS) is a foundational system in China's cybersecurity governance framework. Therefore, accurate analysis and understanding of MLPS requirements are essential. At present, MLPS analysis still relies mainly on manual interpretation of standards and rule-based tools. This makes it hard to provide stable and consistent compliance analysis in complex application scenarios. The rise of large language models has created new opportunities for making MLPS work more intelligent. However, in standards-intensive and security-sensitive scenarios, general-purpose large language models often cannot ensure controllable reasoning or complete understanding of rules. This paper proposes a large language model framework for MLPS that integrates multiple retrieval strategies. It combines hierarchical retrieval, tree-based retrieval, and tokenization-based matching retrieval. This design helps maintain retrieval coverage while reducing the interference of irrelevant context in the reasoning process. To address the requirements of MLPS question answering for clause accuracy, conclusion traceability, and practical deployability, this paper adopts a evaluation method based on multi-dimensional weighted scoring to quantitatively assess model responses. In comparative experiments on ten typical questions, the proposed domain-specific large language model for MLPS achieved higher overall scores.
Problem

Research questions and friction points this paper is trying to address.

MLPS
cybersecurity
large language model
compliance analysis
retrieval fusion
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-path retrieval fusion
domain-specific LLM
hierarchical retrieval
traceable reasoning
multi-dimensional evaluation
πŸ”Ž Similar Papers
Q
Qian Li
The Third Research Institute of Ministry of Public Security, Shanghai 200031, China; Shanghai Engineering Research Center of Cyber and Information Security Evaluation, Shanghai 200031, China
Z
Zhenyan Qi
The Third Research Institute of Ministry of Public Security, Shanghai 200031, China; Shanghai Engineering Research Center of Cyber and Information Security Evaluation, Shanghai 200031, China
L
Liang Shen
The Third Research Institute of Ministry of Public Security, Shanghai 200031, China; Shanghai Engineering Research Center of Cyber and Information Security Evaluation, Shanghai 200031, China
Yuan Zhang
Yuan Zhang
Southeast University
Computer visionMedical image analysis
Y
Yifan Wan
School of Computer Science and Engineering, Southeast University, Nanjing 211189, China
J
Junyuan Ma
School of Cyber Science and Engineering, Southeast University, Nanjing 211189, China
Y
Yining Hu
School of Cyber Science and Engineering, Southeast University, Nanjing 211189, China