Semantic Preprocessing for LLM-based Malware Analysis

📅 2025-06-13
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the weak semantic interpretability of AI models and their disconnection from expert knowledge in malware analysis, this paper proposes a semantics-aware representation method integrating multi-source features and threat intelligence. It unifies static/dynamic analysis outputs, packer detection results, MITRE ATT&CK tactic mappings, and Malware Behavior Catalog (MBC) behavioral modeling into structured JSON reports. This work is the first to jointly embed ATT&CK and MBC knowledge graphs with multimodal analytical features into large language model (LLM) inputs, thereby establishing an organic bridge between domain expertise and LLM interpretability. Evaluated on a realistic, complex dataset, an LLM fine-tuned on these semantic reports achieves a weighted F1-score of 0.94—demonstrating significant improvements over baseline methods in both classification accuracy and analyst comprehensibility.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Interpretability, Analysis, and Evaluation of NLP ModelsKnowledge Representation and Reasoning: Diagnosis and Abductive Reasoning

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Web query analysis, representation and understandingGraph Algorithms and Modeling for the Web: Foundation models and LLMs for Web-related graphs
📝 Abstract
In a context of malware analysis, numerous approaches rely on Artificial Intelligence to handle a large volume of data. However, these techniques focus on data view (images, sequences) and not on an expert's view. Noticing this issue, we propose a preprocessing that focuses on expert knowledge to improve malware semantic analysis and result interpretability. We propose a new preprocessing method which creates JSON reports for Portable Executable files. These reports gather features from both static and behavioral analysis, and incorporate packer signature detection, MITRE ATT&CK and Malware Behavior Catalog (MBC) knowledge. The purpose of this preprocessing is to gather a semantic representation of binary files, understandable by malware analysts, and that can enhance AI models' explainability for malicious files analysis. Using this preprocessing to train a Large Language Model for Malware classification, we achieve a weighted-average F1-score of 0.94 on a complex dataset, representative of market reality.
Problem

Research questions and friction points this paper is trying to address.

Enhancing malware analysis with expert-focused semantic preprocessing
Improving interpretability of AI models in malware detection
Generating JSON reports for Portable Executable files
Innovation

Methods, ideas, or system contributions that make the work stand out.

Semantic preprocessing for expert-focused malware analysis
JSON reports combining static and behavioral features
LLM training with enhanced explainability and classification
🔎 Similar Papers