MASCRDM: Multi-Agent System for Compliance Risk Detection and Mitigation in Training Process of Large Language Models

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitations of static compliance detection and the absence of end-to-end monitoring in large model training by proposing a dynamic compliance intervention mechanism grounded in internal model architectures, thereby transcending conventional input-output filtering paradigms. The proposed method constructs a multi-agent collaborative system that integrates compliance knowledge graphs, specialized large language models (LLMs), and instruction tuning techniques to decompose model nodes and enable real-time risk alerting and mitigation throughout the entire training pipeline. Experimental results demonstrate that this framework effectively reduces discrimination and bias risks while preserving semantic performance, achieving systematic improvements in model compliance.
📝 Abstract
Large Language Models (LLMs) have been applied in various fields. However, ensuring compliance and safety of LLMs, such as avoiding discrimination and bias, still remains a challenge. Current efforts mainly focus on detecting and filtering inputs and outputs of the trained models, rather than studying the intrinsic architecture of the models in real-time. To tackle this challenge, we analyze the LLMs training process and discover two critical issues: 1) Most of the existing methods are predominantly static in their approach to detection and filtering, achieving only localized optimizations without systematically enhancing the compliance of LLMs. 2) Another issue with existing approaches is the lack of real-time risk detection and mitigation across the full training process, which leads to limited flexibility. Motivated by these, we propose MASCRDM (Multi-Agent System for Compliance Risk Detection and Mitigation) during the LLM training process. Firstly, we develop a set of compliance rules based on existing Artificial Intelligence (AI) laws and a compliance-specific LLM with the instruction of compliance law experts. Then, we deconstruct LLMs into several components and identify key nodes based on the compliance knowledge graph. During LLMs training, we implement our multiple agents in the whole process, giving compliance risk alerts and suggestions for LLM developers. Experiments on discrimination and bias benchmark demonstrate that our multi-agent system can effectively improve the compliance while maintaining reasonable semantic performance. The results indicate that our method provides an executable path for mitigating compliance risk from within the LLMs systematically.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Compliance Risk
Training Process
Discrimination and Bias
Risk Detection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multi-Agent System
Compliance Risk Detection
Large Language Models
Training Process
Knowledge Graph
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Yan Zhang
Yan Zhang
Tsinghua University
Computer VisionMulti-Modal Learning
C
Chuming Wei
Academy of Artificial Intelligence and Advanced Technology, Xi’an Jiaotong-Liverpool University
R
Ruien Li
Department of Computer Sciences, University of Wisconsin–Madison
Y
Yaoyao Peng
Law School, University of Chinese Academy of Social Sciences
W
Wusheng Zhang
Department of Computer Science and Technology, Tsinghua University
Guangwen Yang
Guangwen Yang
Professor of Computer Science and Technology, Tsinghua University