evaluate stealth injection attacks

Designs and builds benchmarks, simulation frameworks, and evaluation pipelines for stealthy injection attacks—particularly on persistent-memory/storage—by specifying threat models, injection artifacts, and experiment harnesses that simulate realistic workflows. Defines and implements metrics and instrumentation to measure end-to-end attack success, persistence, stealthiness, detection evasion, coverage of poisoning types (e.g., fact and preference poisoning), and cross-architecture transferability.

evaluatestealthinjectionattacks

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.6
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

StealthCup: Realistic, Multi-Stage, Evasion-Focused CTF for Benchmarking IDS

Nov 21, 2025
MK
Manuel Kern
🏛️ Austrian Institute of Technology | FH Hagenberg | University of Vienna

Existing IDS evaluation methodologies predominantly rely on synthetic datasets or scripted traffic replay, failing to reflect real-world detection capabilities against stealthy, multi-stage APT attacks. This paper introduces the first evasion-oriented, red-team capture-the-flag evaluation framework, conducted in a realistic IT/OT hybrid environment. Professional penetration testers executed 32 MITRE ATT&CK–based attack techniques while synchronously collecting network traffic, host logs, IDS alerts, and post-attack forensic data. Our approach uniquely integrates operational red-team tactics with systematic IDS evaluation, achieving the first high-fidelity replication of national-level APT tradecraft (e.g., Volt Typhoon). Experimental results reveal: (1) 11 attacks evaded all evaluated IDSs; (2) open-source IDSs exhibited false positive rates exceeding 90%, while commercial solutions suffered even higher false negative rates; and (3) all 28 Volt Typhoon techniques were covered, exposing critical blind spots in current IDSs’ ability to detect multi-stage attack chains.

Current approaches under-represent stealthy attacks across both IT and OT domainsEvaluating IDS effectiveness under realistic attack conditions remains challengingExisting benchmarks fail to capture adaptive adversary behavior and multi-stage intrusions

This work addresses a critical gap in existing prompt injection evaluations, which focus solely on final outcomes while overlooking post-attack trajectory divergence and degradation of authorized functionalities. To remedy this, the authors introduce ContainmentBench, a novel benchmark that pioneers a trajectory-based logging framework coupled with phased utility assessment. Evaluated within a sandboxed environment, this approach enables fine-grained analysis of tool-using LLM agents across four dimensions: trajectory propagation, policy compliance, recovery capability, and task completion. By leveraging contamination markers, structured authorization logs, and multidimensional metrics, the framework uncovers hidden containment disparities even when terminal outputs appear identical. Across 17,640 rollback experiments, 73.5% exhibited significant differences in either trajectory or utility. Moreover, the proposed trusted ledger and strong tool-boundary policies elevate authorized task completion rates to 0.8567 and 0.9233, respectively.

post-injection containmentprompt injectionsecurity benchmark

This paper addresses the lack of standardized, rigorous evaluation benchmarks for large language models (LLMs) in adversarial cybersecurity applications—particularly penetration testing. We systematically review 15 prototype systems and their empirical testing practices drawn from 16 scholarly works. Through comparative analysis of testbed architectures, evaluation metric frameworks, and security experimentation methodologies, we identify three pervasive deficiencies: poor reproducibility, weak real-world representativeness (especially the limited transferability of CTF-based scenarios to actual red-teaming), and narrow assessment scope (lacking established baselines, qualitative analysis, and human-factor considerations). To address these gaps, we propose the first dedicated evaluation framework for LLM-driven adversarial security, featuring an extensible testbed, standardized baseline construction, a hybrid quantitative–qualitative metric suite, and human-in-the-loop assessment dimensions. Based on this framework, we derive seven actionable best-practice recommendations—establishing a scientifically grounded, robust, and operationally viable evaluation paradigm for the field.

Assessing benchmarking practices in LLM-based penetration testingBridging gaps between security research and real-world practiceEvaluating LLM-driven offensive cybersecurity tools' effectiveness

This work addresses the limitations of existing adversarial simulation tools, which rely on agent-based instrumentation of target systems, often leaving anomalous artifacts and failing to faithfully replicate human attacker behavior—particularly in critical phases of the cyber kill chain such as initial access and interactive operations. To overcome these shortcomings, the authors propose and implement an open-source attack scripting language coupled with an agentless execution engine that closely emulates real-world attacker tactics. This approach enables high-fidelity, interactive simulation of complete kill chain stages, including initial access, privilege escalation, and lateral movement. Experimental results demonstrate that system logs generated by this method exhibit significantly greater behavioral similarity to those produced by actual human-driven attacks, thereby enhancing the realism and effectiveness of security testing and intrusion detection research.

adversary emulationattack automationcyber attack scenarios

Formalizing and Benchmarking Prompt Injection Attacks and Defenses

Oct 19, 2023
YL
Yupei Liu
🏛️ The Pennsylvania State University | Duke University

Existing research lacks systematic modeling and standardized evaluation of prompt injection attacks and defenses in LLM-based integrated applications. Method: We propose the first general formal attack framework that unifies five known attack categories and enables derivation of novel composite attacks; we further construct the first open-source, cross-model (e.g., GPT, Llama, Claude) and cross-task (e.g., QA, summarization, reasoning—seven tasks total) benchmark, covering ten defense mechanisms. Contribution/Results: Through red-team/blue-team adversarial experiments, we demonstrate that most existing defenses fail under complex, realistic scenarios. We release Open-Prompt-Injection—a reproducible, multidimensional quantitative evaluation platform—to advance standardization and community collaboration in prompt security research.

Benchmarking existing attacks and defenses across multiple modelsEstablishing common evaluation framework for future research developmentFormalizing prompt injection attacks in LLM applications systematically

Latest Papers

What's happening recently
View more

This work addresses latent security risks introduced by persistent components—such as memory and tools—in agent frameworks, where malicious payloads can remain dormant across tasks and trigger harm in subsequent benign requests. The study presents the first systematic modeling of such cross-component delayed risks, introducing the concept of a “persistence risk lifecycle” and constructing a benchmark comprising 328 executable cases spanning seven carrier types. Employing a multi-stage evaluation method based on execution traces, the framework tracks attacks from initial injection through persistence to eventual policy violation. Experimental results reveal that defensive efficacy is highly dependent on both carrier type and agent–model configuration, and that relying solely on end-to-end attack success rates obscures critical differences in risk propagation dynamics, thereby offering fine-grained insights for secure agent system design.

agent harnessesattack propagationdelayed security

This work addresses the lack of transparency in autonomous penetration testing agents when verifying vulnerabilities under deceptive responses, where conflicting evidence handling and decision logic are difficult to trace. To this end, the paper introduces ATOBench, an evaluation framework that enables the first observable verification chain by injecting registered response transformations at runtime, aligning original and transformed test snippets, and reconstructing source links to track actions, evidence recovery, termination decisions, and report justification. The framework formalizes three frozen observation contracts—exploit proof, resource ownership, and reusable artifacts—to structurally assess evidence processing. Evaluation across 450 test snippets on five model pipelines reveals that high activity levels can obscure verification chain breaks, while successful recovery hinges on the discovery and retention of critical evidence, demonstrating ATOBench’s effectiveness in exposing agent verification behavior under untrusted observations.

agent evaluationautonomous penetration testingdeceptive responses

Hot Scholars

SJ

Shouling Ji

Professor, Zhejiang University & Georgia Institute of Technology
Data-driven SecurityAI SecuritySoftware ScurityPrivacy
CZ

Chunyi Zhou

Zhejiang University
Cyberspace SecurityMachine Learning PrivacyFederated Learning
TD

Tianyu Du

Zhejiang University
AI SecurityAdversarial Machine Learning
JJ

Jinghan Jia

Michigan State University
Machine LearningGenerative AIAI SafetyEfficient AI.
JL

Junxian Li

NSEC lab,Shanghai Jiaotong University
AI securityReasoningData Mining