pikit: A Composable Toolkit for Indirect Prompt Injection Research and Evaluation

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of systematic tools for evaluating indirect prompt injection threats by proposing a composable assessment framework spanning three dimensions: attacks, channels, and defenses. Built upon a decorator-registry architecture, the framework enables non-invasive extensibility and flexibly combines arbitrary attacks with channels through a unified craft() interface, while integrating an LLM agent testing environment. Experimental results demonstrate that the proposed defense strategies reduce the success rate of high-risk attacks by 71.8%, and that offline detection exhibits high-precision characteristics. By providing a standardized and scalable solution, this work advances the security evaluation of indirect prompt injection in large language model applications.
📝 Abstract
Indirect prompt injection embeds malicious instructions within external content retrieved by LLM-based agents, altering target behavior without user authorization. We introduce pikit, a research toolkit designed to systematically evaluate these threats across three core dimensions: attacks (13 methods), channels (16 carriers across text and file modes), and defenses (9 prevention strategies and 3 offline detection baselines). Built on a decorator-based registry, pikit enables seamless extension of custom components without modifying core code, while a unified craft() API composes arbitrary attacks and channels in a single call. We evaluated the toolkit on the pi coding agent powered by an anonymized LLM in a production-like environment. Benchmarking 9 prevention strategies against high-risk attacks yields a 71.8\% relative reduction in attack success rate, with few\_shot\_warning and instruction\_hierarchy providing the strongest protection. Offline detection baselines achieve perfect precision but low recall, demonstrating that heuristic detectors complement rather than replace prompt-level defenses. To ensure reproducibility, each run automatically logs full prompts, agent event traces, session transcripts, and verdict records. Our code is available at https://github.com/Tencent/AI-Infra-Guard/tree/main/Research/pikit.
Problem

Research questions and friction points this paper is trying to address.

Indirect Prompt Injection
LLM-based Agents
Security Evaluation
Adversarial Attacks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Indirect Prompt Injection
Composable Toolkit
Decorator-based Registry
LLM Agent Security
Defense Evaluation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Zonghao Ying
Zonghao Ying
SKLCCSE, BUAA
Trustworthy AI
X
Xiangfan Wu
Tencent Zhuque Lab
B
Bo Yang
Tencent Zhuque Lab
H
Huiyu Wu
Tencent Zhuque Lab
Xing Zheng
Xing Zheng
Ph.D. of University of California, Riverside
Sensor fusionSLAMVIO
H
Huangsheng Cheng
Tencent Zhuque Lab
X
Xiaorong Shi
Tencent Zhuque Lab
J
Jing Guo
Tencent Zhuque Lab