Scalable APT Malware Classification via Parallel Feature Extraction and GPU-Accelerated Learning

📅 2025-04-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Fine-grained classification of malicious executables associated with Advanced Persistent Threat (APT) groups remains challenging. Method: We propose an end-to-end, automated classification framework based on assembly opcodes. It introduces a novel opcode-level parallel reverse-engineering pipeline using Radare2 and multiprocessing for efficient feature extraction. Departing from metadata-dependent n-gram models, we design a GPU-accelerated CNN architecture (PyTorch + CUDA) that takes raw instruction sequences as input, and rigorously benchmark against SVM, KNN, and decision tree baselines. Contribution/Results: This is the first work to achieve high-throughput, low-latency APT family attribution at the opcode level. Evaluated on a real-world APT dataset, our method achieves 98.3% mean accuracy, accelerates inference 17× over CPU-based baselines, and supports real-time processing of over 1,000 samples per minute.

Technology Category

Machine Learning: Hardware-aware MLComputer Vision: Adversarial Attacks & RobustnessNatural Language Processing: Sentence-level Semantics, Textual Inference, etc.

Application Category

Systems and Infrastructure for Web, Mobile and WoT: Applied ML and AI for Web-based mobile applicationsUser Modeling, Personalization and Recommendation: Attacks and countermeasures in recommendation systemsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
📝 Abstract
This paper presents an underlying framework for both automating and accelerating malware classification, more specifically, mapping malicious executables to known Advanced Persistent Threat (APT) groups. The main feature of this analysis is the assembly-level instructions present in executables which are also known as opcodes. The collection of such opcodes on many malicious samples is a lengthy process; hence, open-source reverse engineering tools are used in tandem with scripts that leverage parallel computing to analyze multiple files at once. Traditional and deep learning models are applied to create models capable of classifying malware samples. One-gram and two-gram datasets are constructed and used to train models such as SVM, KNN, and Decision Tree; however, they struggle to provide adequate results without relying on metadata to support n-gram sequences. The computational limitations of such models are overcome with convolutional neural networks (CNNs) and heavily accelerated using graphical compute unit (GPU) resources.
Problem

Research questions and friction points this paper is trying to address.

Automating and accelerating APT malware classification
Overcoming computational limits in opcode-based feature extraction
Improving classification accuracy without relying on metadata
Innovation

Methods, ideas, or system contributions that make the work stand out.

Parallel computing for feature extraction
GPU-accelerated CNN for malware classification
N-gram datasets with assembly-level opcodes
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.