PatchHolmes: Agentic Patch Retrieval via Listwise Selection

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of vulnerability management caused by the widespread absence of patch links in CVE records by proposing a two-stage retrieval framework. First, a hybrid retriever performs preliminary filtering over commits in local Git repositories. Subsequently, a large language model (LLM) agent is introduced to overcome the limitations of conventional pointwise scoring, precisely identifying remediation patches through a listwise selection mechanism within a single conversational turn. Experimental results demonstrate that, without fine-tuning, the proposed approach significantly outperforms existing baselines in Recall@1, validating the core efficacy of the listwise agent mechanism for accurate patch retrieval.
📝 Abstract
Patch retrieval, the task of finding the commit that fixes a known vulnerability, is the foundation of vulnerability management workflows, yet 60% to 63% of CVEs in the major advisory databases lack a patch link. We present PatchHolmes, a two-phase patch retrieval system that pairs a hybrid first-stage retriever with an agentic second-stage inspection loop. Unlike pointwise prior work that scores each candidate independently, the Phase 2 agent reads the top-100 listwise: it sees the full candidate list at once and selectively reads 3 to 10 commits through four budgeted tools before submitting a single best commit. On GitHubAD, PatchHolmes beats the pointwise binary classifier Favia by 25.34% Recall@1 and the retrieve-and-CoT baseline IRCoT by 31.40%, at one agent conversation per CVE versus Favia's ten; with the candidate set held identical, the agent adds 27.32% Recall@1 over taking the retriever's top candidate, and the same agent, transferred unchanged to PatchFinder_top10, lifts Recall@1 from PatchFinder's own top-1 pick (24.28%) to 39.86%. Swapping the LLM backbone within the Qwen family changes Recall@1 by under 1%, and a second model family (gpt-oss) stays far above the no-agent floor, so the gain comes from the listwise agent loop; the entire system runs on a frozen open-weight model over a local Git repository, without fine-tuning or external search APIs.
Problem

Research questions and friction points this paper is trying to address.

Patch Retrieval
Vulnerability Management
CVE
Commit Matching
Security Advisory
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic Patch Retrieval
Listwise Selection
Vulnerability Management
Budgeted Tool Use
Zero-shot Transfer
🔎 Similar Papers
No similar papers found.