AutoDP-LLM: automating data pre-processing for intrusion detection systems using large language models

📅 2026-10-01
🏛️ Journal of Supercomputing
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the reliance on manual trial-and-error and high computational overhead in data preprocessing for intrusion detection systems (IDS) by proposing a large language model (LLM)-based framework for automated pipeline generation. The approach integrates LLM agents with deterministic planning to optimize feature selection through the synergy of semantic reasoning and statistical evidence. It enables adaptive feature retention without requiring predefined feature budgets and supports parallelized code synthesis. Experimental results demonstrate that the proposed framework achieves competitive detection performance on benchmark datasets while significantly reducing manual effort, efficiently generating compact and executable preprocessing pipelines.
📝 Abstract
The increasing complexity and scale of modern cyber-attacks demand intelligent and computationally efficient Intrusion Detection Systems (IDS). However, designing effective data pre-processing pipelines traditionally involves substantial trial-and-error effort and repeated evaluation of alternative configurations. For large, high-dimensional network traffic data, this process can create a significant computational burden. In this work, we propose AutoDP-LLM, an automated pre-processing framework designed to reduce manual pipeline development and computational overhead. Specifically, AutoDP-LLM leverages Large Language Models (LLMs) to autonomously generate and validate executable data pre-processing pipelines. The framework combines deterministic host-side planning with LLM-based specialist agents to formulate data-processing strategies, synthesize executable code, and adaptively determine retained feature sets using semantic reasoning and training-derived statistical evidence, without requiring a predefined feature budget. Focusing on multiclass intrusion detection, we evaluate AutoDP-LLM on the UNSW-NB15 and NSL-KDD benchmark datasets using multiple downstream classifiers. Comparative experiments against conventional feature-selection methods show that AutoDP-LLM achieves competitive detection performance while automating the generation of compact and executable pre-processing pipelines. Component-level ablation experiments further demonstrate the complementary contributions of the semantic and statistical feature-reduction components. The repeated generation, validation, execution, and assessment of candidate pipelines are amenable to parallel execution, highlighting the potential of scalable computing environments, including high-performance computing (HPC) systems, to support automated IDS pipeline development.
Problem

Research questions and friction points this paper is trying to address.

Intrusion Detection Systems
Data Pre-processing
Large Language Models
Feature Selection
Computational Overhead
Innovation

Methods, ideas, or system contributions that make the work stand out.

Large Language Models
Automated Data Pre-processing
Intrusion Detection Systems
Feature Selection
High-Performance Computing
🔎 Similar Papers
B
Bao-Phong Nguyen
Business AI Lab, College of Technology, National Economics University, Hanoi, Vietnam
G
Gia-Khanh Pham
Business AI Lab, College of Technology, National Economics University, Hanoi, Vietnam
T
Thai-Duong Do
Business AI Lab, College of Technology, National Economics University, Hanoi, Vietnam
Mai Xuan Trang
Mai Xuan Trang
Phenikaa AIoT Lab, Phenikaa University
Machine LearningDeep LearningCloud computingServices computing
Minh-Tuan Le
Minh-Tuan Le
Posts and Telecommunications Institute of Technology
MIMOSTBCSpatial ModulationMassive MIMO
X
Xuan-Nam Tran
Advanced Wireless Communications Group, Le Quy Don Technical University, Hanoi, Vietnam
H
Huan Vu
Business AI Lab, College of Technology, National Economics University, Hanoi, Vietnam
T
Tien-Cuong Nguyen
VNPT AI, VNPT Group, Hanoi, Vietnam
Vu-Duc Ngo
Vu-Duc Ngo
MobiFone R&D Center
PHY designSoC and NoCAI&ML for 6G
Thien Van Luong
Thien Van Luong
Business AI Lab, National Economics University, Vietnam
Medical AIfraud detectiontime-serieswireless communications