🤖 AI Summary
This study investigates the practical efficacy of GitHub Copilot’s code review capability in detecting security vulnerabilities. Method: We constructed a manually annotated, multilingual dataset of open-source project vulnerabilities—covering SQL injection, cross-site scripting (XSS), and insecure deserialization—and conducted controlled experiments to evaluate Copilot’s feedback along three dimensions: accuracy, vulnerability coverage, and risk-level alignment. Contribution/Results: Empirical results show that Copilot detects fewer than 5% of high-severity vulnerabilities, predominantly flagging low-risk stylistic issues instead. Its security detection performance is substantially inferior to both specialized static application security testing (SAST) tools and human security audits. To our knowledge, this is the first empirical study to expose fundamental limitations of AI-powered programming assistants in security-critical code review tasks. Our findings challenge the prevailing assumption that AI can substitute for traditional security auditing and provide critical evidence for defining the security boundaries of AI-assisted development and designing effective human–AI collaboration paradigms.
📝 Abstract
As software development practices increasingly adopt AI-powered tools, ensuring that such tools can support secure coding has become critical. This study evaluates the effectiveness of GitHub Copilot's recently introduced code review feature in detecting security vulnerabilities. Using a curated set of labeled vulnerable code samples drawn from diverse open-source projects spanning multiple programming languages and application domains, we systematically assessed Copilot's ability to identify and provide feedback on common security flaws. Contrary to expectations, our results reveal that Copilot's code review frequently fails to detect critical vulnerabilities such as SQL injection, cross-site scripting (XSS), and insecure deserialization. Instead, its feedback primarily addresses low-severity issues, such as coding style and typographical errors. These findings expose a significant gap between the perceived capabilities of AI-assisted code review and its actual effectiveness in supporting secure development practices. Our results highlight the continued necessity of dedicated security tools and manual code audits to ensure robust software security.