Adaptive Code Revision Attacks on AI Pull Request Reviewers

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of AI-assisted code review feedback to adversarial exploitation, wherein attackers may circumvent security detection by patching superficial flaws while preserving deeper risks. To investigate this threat, we propose Adaptive Feedback-Guided Code Revision Attacks (AFCRA), which integrates large language model interactions, automated code analysis, and adversarial prompt engineering, alongside a dedicated benchmark for validation. This work is the first to demonstrate that adversaries can dynamically modify code to conceal vulnerabilities by exploiting review feedback, thereby transcending traditional text-based persuasion paradigms. Experimental results indicate that AFCRA increases attack success rates by 2.5× and 12.5× on Sonnet 5 and GPT-5.5, respectively. These findings expose a novel attack surface in AI-driven development pipelines and offer actionable recommendations for mitigation.
📝 Abstract
Pull-request review protects software before new code reaches users, helping prevent vulnerabilities that could expose users to attacks. AI agents increasingly perform these reviews and explain which problems need fixing. However, for an attacker submitting vulnerable code, this feedback also reveals what changes may secure approval. Existing PR attacks seek such approval through persuasive text and comments while keeping executable code fixed. This leaves unclear whether an attacker can use the feedback to repair the reported problem while preserving a vulnerability in the revised code. We therefore conduct an empirical study of this threat using AFCRA (Adaptive Feedback-guided Code Revision Attack). To distinguish successful attacks from genuine repairs, we construct AFCRA-Bench from 159 disclosed vulnerabilities, with executable exploits to verify vulnerabilities in code. Across five-round interactions with Sonnet 5 and GPT-5.5 reviewers, AFCRA reaches success rates 2.5x and 12.5x those of the strongest evaluated text- or comment-based attack. Case studies of these successes show how reviewers accept repairs of reported problems while overlooking surviving vulnerabilities. These findings establish feedback-guided code revision as a threat to automated PR review. To address this threat, we derive actionable implications for researchers, AI providers, PR reviewers, and PR authors on securing AI-assisted development.
Problem

Research questions and friction points this paper is trying to address.

Pull Request Review
Code Revision Attack
AI Security
Vulnerability Preservation
Feedback-guided Attack
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive Code Revision Attack
AI Pull Request Reviewer
Feedback-guided Attack
Vulnerability Preservation
AFCRA-Bench
🔎 Similar Papers
No similar papers found.