Closing the Gap: A User Study on the Real-world Usefulness of AI-powered Vulnerability Detection & Repair in the IDE

📅 2026-04-11
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
AI-based vulnerability detection and repair tools for professional developers suffer from high false-positive rates, inapplicable patch suggestions, and poor integration into development workflows. This paper presents the first industrial-scale empirical study deploying DeepVulGuard—an AI-powered tool integrating state-of-the-art vulnerability detection/repair models, context-aware scanning, natural language generation (NLG) for explanatory feedback, and IDE-embedded conversational interaction—across 24 real-world projects totaling over 1.7 million lines of code. We systematically evaluate its practicality, trustworthiness, and workflow integration. We propose a novel human-AI collaboration paradigm centered on “explanation–confidence feedback–interactive co-refinement.” Empirical evaluation yielded 170 vulnerability alerts and 50 repair suggestions, uncovering critical bottlenecks and yielding actionable deployment optimization strategies. Our findings provide empirical foundations and methodological guidance for the trustworthy adoption of AI coding assistants in professional software engineering contexts.

Technology Category

Humans and AI: Human-AI Collaboration / Human-AI TeamingComputer Vision: Adversarial Attacks & RobustnessNatural Language Processing: Safety and Robustness

Application Category

Economics, Online Markets and Human Computation: Trust and reliance of crowd workers and data experts on GenAIResponsible Web: Human-perceived consequences of algorithmic deployment on the webSecurity and Privacy: Large-scale security measurements
📝 Abstract
This paper presents the first empirical study of a vulnerability detection and fix tool with professional software developers on real projects that they own. We implemented DeepVulGuard, an IDE-integrated tool based on state-of-the-art detection and fix models, and show that it has promising performance on benchmarks of historic vulnerability data. DeepVulGuard scans code for vulnerabilities (including identifying the vulnerability type and vulnerable region of code), suggests fixes, provides natural-language explanations for alerts and fixes, leveraging chat interfaces. We recruited 17 professional software developers at Microsoft, observed their usage of the tool on their code, and conducted interviews to assess the tool's usefulness, speed, trust, relevance, and workflow integration. We also gathered detailed qualitative feedback on users' perceptions and their desired features. Study participants scanned a total of 24 projects, 6.9k files, and over 1.7 million lines of source code, and generated 170 alerts and 50 fix suggestions. We find that although state-of-the-art AI-powered detection and fix tools show promise, they are not yet practical for real-world use due to a high rate of false positives and non-applicable fixes. User feedback reveals several actionable pain points, ranging from incomplete context to lack of customization for the user's codebase. Additionally, we explore how AI features, including confidence scores, explanations, and chat interaction, can apply to vulnerability detection and fixing. Based on these insights, we offer practical recommendations for evaluating and deploying AI detection and fix models. Our code and data are available at https://doi.org/10.6084/m9.figshare.26367139.
Problem

Research questions and friction points this paper is trying to address.

Smart Tools
Code Vulnerabilities
Professional Programmers
Innovation

Methods, ideas, or system contributions that make the work stand out.

DeepVulGuard
AI-assisted programming
Automatic vulnerability detection and repair
🔎 Similar Papers
No similar papers found.