MASCRDM: Multi-Agent System for Compliance Risk Detection and Mitigation in Training Process of Large Language Models
This study addresses the limitations of static compliance detection and the absence of end-to-end monitoring in large model training by proposing a dynamic compliance intervention mechanism grounded in internal model architectures, thereby transcending conventional input-output filtering paradigms. The proposed method constructs a multi-agent collaborative system that integrates compliance knowledge graphs, specialized large language models (LLMs), and instruction tuning techniques to decompose model nodes and enable real-time risk alerting and mitigation throughout the entire training pipeline. Experimental results demonstrate that this framework effectively reduces discrimination and bias risks while preserving semantic performance, achieving systematic improvements in model compliance.