CAPTURE: Context-Aware Prompt Injection Testing and Robustness Enhancement

📅 2025-05-18
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the dual challenges of insufficient contextual awareness and over-defensiveness in large language model (LLM) prompt injection detection—manifesting as high false-negative rates under adversarial conditions and high false-positive rates on benign inputs—this paper introduces the first lightweight benchmark jointly evaluating attack detection capability and over-defense propensity. Methodologically, we propose a context-aware evaluation framework integrating dynamic context modeling, adversarial sample generation, and fine-grained defense behavior analysis, augmented by a minimally supervised bias calibration mechanism enabling robust assessment with only a few domain-specific examples. Our contributions are threefold: (1) the first simultaneous quantification of both false-negative and false-positive rates, overcoming limitations of static testing paradigms; (2) empirical identification of critical vulnerabilities in state-of-the-art protective models; and (3) a reproducible, generalizable, and trustworthy evaluation standard for LLM security.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Safety and RobustnessComputer Vision: Adversarial Attacks & Robustness

Application Category

User Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendationSearch and Retrieval-Augmented AI: Search Tool Learning with LLM: Teaching LLMs to invoke search and make use of retrieved informationSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
📝 Abstract
Prompt injection remains a major security risk for large language models. However, the efficacy of existing guardrail models in context-aware settings remains underexplored, as they often rely on static attack benchmarks. Additionally, they have over-defense tendencies. We introduce CAPTURE, a novel context-aware benchmark assessing both attack detection and over-defense tendencies with minimal in-domain examples. Our experiments reveal that current prompt injection guardrail models suffer from high false negatives in adversarial cases and excessive false positives in benign scenarios, highlighting critical limitations.
Problem

Research questions and friction points this paper is trying to address.

Assessing prompt injection risks in context-aware settings
Evaluating over-defense tendencies in guardrail models
Improving detection accuracy with minimal in-domain examples
Innovation

Methods, ideas, or system contributions that make the work stand out.

Context-aware benchmark for prompt injection testing
Minimal in-domain examples for robustness enhancement
Detects high false negatives and excessive false positives
💼 Related Jobs
No related jobs found.
Pure Storage