GRAML: Graph-Grounded Reasoning and Multi-Task Learning for LLM-Based Software Vulnerability Detection

📅 2026-10-07
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the performance bottleneck of large language models in software vulnerability detection caused by the absence of control-flow and data-flow information. To overcome this limitation, we propose a unified framework integrating graph evidence, vulnerability description generation, and multi-task learning. Methodologically, static analysis is employed to extract code property graphs as structural evidence, guiding GPT-5 to perform graph-driven Tree-of-Thought Vulnerability Reasoning (ToT-VR), while a four-task joint training dataset is constructed. Experimental results demonstrate that the proposed approach achieves average F1 scores of 66.67%–68.70%, outperforming state-of-the-art baselines by up to 30.92%. It significantly surpasses standard Chain-of-Thought prompting and naive graph serialization schemes, effectively enhancing the reliability of vulnerability detection.
📝 Abstract
Large Language Models (LLMs) have been widely applied to software vulnerability detection. However, their performance is often limited by insufficient use of control-flow and data-flow information. In this paper, we propose GRAML, a framework that combines graph evidence, vulnerability description generation, and multi-task training. GRAML first performs static analysis on C/C++ programs to extract critical source lines and typed line relations as structural evidence. It then uses this evidence to guide GPT-5 through the Tree-of-Thought-guided Vulnerability Reasoning (ToT-VR) process and generate vulnerability descriptions. These descriptions are further combined with Detection, Localization, and Assessment samples to build a unified four-task training dataset. We evaluate GRAML on an in-distribution (ID) test set and six out-of-distribution (OOD) datasets. The results show that GRAML achieves average F1 scores ranging from 66.67% to 68.70%, outperforming state-of-the-art baselines by up to 30.92%. Ablation experiments further show that ToT-VR and graph-guided vulnerability descriptions improve detection performance compared with standard Chain-of-Thought (CoT) reasoning and raw Code Property Graph (CPG) serializations. These findings provide practical guidance for building more reliable and secure software engineering systems with large language models.
Problem

Research questions and friction points this paper is trying to address.

Software Vulnerability Detection
Large Language Models
Control-flow and Data-flow Information
Innovation

Methods, ideas, or system contributions that make the work stand out.

Graph-Grounded Reasoning
Multi-Task Learning
Tree-of-Thought
Vulnerability Detection
Large Language Models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Ruohan Li
Qingdao University of Technology, Qingdao, Shandong, China
M
Miaoqing Tian
Qingdao University of Technology, Qingdao, Shandong, China
H
Haipeng Qu
Ocean University of China, Qingdao, Shandong, China
Ruobing Jiang
Ruobing Jiang
PhD student of Computer Science and Engineering, Shanghai Jiao Tong University
vehicular ad hoc networkswireless sensor networkscompressive sensingintelligent transportation system