I Can't Believe It's Not a Valid Exploit

📅 2026-02-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the tendency of existing evaluation methodologies to overestimate the effectiveness of large language models (LLMs) in generating valid proof-of-concept (PoC) exploits for software vulnerabilities. To this end, we propose PoC-Gym, a novel framework that systematically integrates static analysis tools to guide LLMs—including Claude Sonnet 4, GPT-5 Medium, and gpt-oss-20b—in generating Java vulnerability PoCs, followed by rigorous validation through both automated and manual review. Our experiments demonstrate that static analysis guidance improves PoC generation success rates by 21%; however, manual inspection reveals that 71.5% of the generated PoCs are in fact invalid. This discrepancy exposes a significant bias in current evaluation practices and underscores the critical contribution of our approach in enhancing assessment rigor and methodological innovation.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Code Generation / Program Synthesis from Natural LanguageSearch and Optimization: Evaluation and Analysis

Application Category

Economics, Online Markets and Human Computation: Cost models of using LLMs in production systemsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSearch and Retrieval-Augmented AI: Web evaluation methodologies and metrics
📝 Abstract
Recently Large Language Models (LLMs) have been used in security vulnerability detection tasks including generating proof-of-concept (PoC) exploits. A PoC exploit is a program used to demonstrate how a vulnerability can be exploited. Several approaches suggest that supporting LLMs with additional guidance can improve PoC generation outcomes, motivating further evaluation of their effectiveness. In this work, we develop PoC-Gym, a framework for PoC generation for Java security vulnerabilities via LLMs and systematic validation of generated exploits. Using PoC-Gym, we evaluate whether the guidance from static analysis tools improves the PoC generation success rate and manually inspect the resulting PoCs. Our results from running PoC-Gym with Claude Sonnet 4, GPT-5 Medium, and gpt-oss-20b show that using static analysis for guidance and criteria lead to 21% higher success rates than the prior baseline, FaultLine. However, manual inspection of both successful and failed PoCs reveals that 71.5% of the PoCs are invalid. These results show that the reported success of LLM-based PoC generation can be significantly misleading, which is hard to detect with current validation mechanisms.
Problem

Research questions and friction points this paper is trying to address.

PoC generation
LLM-based exploit
validation
security vulnerability
static analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

PoC generation
Large Language Models
static analysis
vulnerability exploitation
systematic validation
🔎 Similar Papers
No similar papers found.