🤖 AI Summary
To address the weak semantic interpretability of AI models and their disconnection from expert knowledge in malware analysis, this paper proposes a semantics-aware representation method integrating multi-source features and threat intelligence. It unifies static/dynamic analysis outputs, packer detection results, MITRE ATT&CK tactic mappings, and Malware Behavior Catalog (MBC) behavioral modeling into structured JSON reports. This work is the first to jointly embed ATT&CK and MBC knowledge graphs with multimodal analytical features into large language model (LLM) inputs, thereby establishing an organic bridge between domain expertise and LLM interpretability. Evaluated on a realistic, complex dataset, an LLM fine-tuned on these semantic reports achieves a weighted F1-score of 0.94—demonstrating significant improvements over baseline methods in both classification accuracy and analyst comprehensibility.
📝 Abstract
In a context of malware analysis, numerous approaches rely on Artificial Intelligence to handle a large volume of data. However, these techniques focus on data view (images, sequences) and not on an expert's view. Noticing this issue, we propose a preprocessing that focuses on expert knowledge to improve malware semantic analysis and result interpretability. We propose a new preprocessing method which creates JSON reports for Portable Executable files. These reports gather features from both static and behavioral analysis, and incorporate packer signature detection, MITRE ATT&CK and Malware Behavior Catalog (MBC) knowledge. The purpose of this preprocessing is to gather a semantic representation of binary files, understandable by malware analysts, and that can enhance AI models' explainability for malicious files analysis. Using this preprocessing to train a Large Language Model for Malware classification, we achieve a weighted-average F1-score of 0.94 on a complex dataset, representative of market reality.