ENDOPROMPT: Victim-Side Pseudo-References for Utility Degradation

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation that existing prompt injection attacks predominantly rely on harmful content, manual annotations, or predefined targets, thereby failing to purely expose the utility vulnerabilities of large language models. To this end, this work proposes a victim-side pseudo-reference supervision mechanism that operates without benchmark feedback. Specifically, it leverages clean continuations of unlabeled instructions as pseudo-references, integrating local search with preference optimization to train a utility-degrading prefix generator. This research pioneers a pure utility attack paradigm devoid of harmful content, achieving generative prefix learning through white-box search and reward refinement. Extensive evaluations across four models and seven benchmarks demonstrate an average utility degradation of 26.8 percentage points, with significant performance deterioration observed in 27 out of 28 metrics, effectively revealing critical utility weaknesses in large language models.
📝 Abstract
Prompt injection can degrade benign task performance without eliciting harmful content. Yet many attack objectives depend on task labels or predefined target responses. We present ENDOPROMPT, a white-box method that learns utility-degrading prefixes from unlabeled instructions. Its generator takes the request text as input. Clean victim continuations serve as pseudo-references: local search identifies prefixes that reduce continuation likelihood, and preference fitting on comparisons within the same instruction, followed by reward refinement, distills this signal into a generator. At deployment, the generator produces one prefix per request without further victim-side search. Across four instruction-tuned models and the complete splits of seven benign benchmarks, ENDOPROMPT yields a mean utility change of -26.8 percentage points; 27 of 28 cells are negative. Failure analysis reveals output expansion and prefix reuse; the controls do not establish a degradation advantage from request matching. Victim-derived supervision can reveal utility weaknesses without benchmark feedback or prescribed failure responses. The code will be released upon acceptance.
Problem

Research questions and friction points this paper is trying to address.

prompt injection
utility degradation
unlabeled instructions
white-box attack
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

Prompt Injection
Utility Degradation
Pseudo-References
Preference Fitting
White-box Attack
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Q
Qingyu Wu
Defense Innovation Institute, Academy of Military Science, Beijing, China
Zeyu Feng
Zeyu Feng
The University of Sydney
Y
Yongda Yu
Nanjing University, Nanjing, China
Y
Yuzhe Luo
Defense Innovation Institute, Academy of Military Science, Beijing, China
Hua Cheng
Hua Cheng
Associate Professor in School of Physics, Nankai University
Metamaterials in optics and acoustics