Detecting Cybersecurity Threats by Integrating Explainable AI with SHAP Interpretability and Strategic Data Sampling

📅 2026-02-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work proposes an end-to-end trustworthy detection framework to address three major challenges in cybersecurity threat detection: large-scale data volume, high risk of feature leakage, and opaque model decisions. The framework uniquely integrates strategic sampling—preserving class distribution to enhance training efficiency—an automated data leakage prevention mechanism, and model-agnostic SHAP-based interpretability analysis. Experimental evaluation on the CIC-IDS2017 dataset demonstrates that the proposed approach significantly reduces computational overhead while maintaining high detection performance. Furthermore, it delivers actionable explanations for security analysts, thereby facilitating the practical deployment of trustworthy AI in Security Operations Centers (SOCs).

Technology Category

Computer Vision: Interpretability, Explainability, and TransparencyMachine Learning: Transparent, Interpretable, Explainable MLNatural Language Processing: Interpretability, Analysis, and Evaluation of NLP Models

Application Category

Security and Privacy: Data transparency and provenanceWeb Mining and Content Analysis: Large pretrained models with web dataUser Modeling, Personalization and Recommendation: Attacks and countermeasures in recommendation systems
📝 Abstract
The critical need for transparent and trustworthy machine learning in cybersecurity operations drives the development of this integrated Explainable AI (XAI) framework. Our methodology addresses three fundamental challenges in deploying AI for threat detection: handling massive datasets through Strategic Sampling Methodology that preserves class distributions while enabling efficient model development; ensuring experimental rigor via Automated Data Leakage Prevention that systematically identifies and removes contaminated features; and providing operational transparency through Integrated XAI Implementation using SHAP analysis for model-agnostic interpretability across algorithms. Applied to the CIC-IDS2017 dataset, our approach maintains detection efficacy while reducing computational overhead and delivering actionable explanations for security analysts. The framework demonstrates that explainability, computational efficiency, and experimental integrity can be simultaneously achieved, providing a robust foundation for deploying trustworthy AI systems in security operations centers where decision transparency is paramount.
Problem

Research questions and friction points this paper is trying to address.

Cybersecurity Threat Detection
Explainable AI
SHAP Interpretability
Strategic Data Sampling
Data Leakage Prevention
Innovation

Methods, ideas, or system contributions that make the work stand out.

Explainable AI
SHAP
Strategic Sampling
Data Leakage Prevention
Cybersecurity Threat Detection
N
Norrakith Srisumrith
Dept. of Digital Network and Information Security Management, King Mongkut’s University of Technology North Bangkok, Bangkok, Thailand
S
Sunantha Sodsee
Dept. of Digital Network and Information Security Management, King Mongkut’s University of Technology North Bangkok, Bangkok, Thailand