🤖 AI Summary
This study addresses the limitation that existing prompt injection attacks predominantly rely on harmful content, manual annotations, or predefined targets, thereby failing to purely expose the utility vulnerabilities of large language models. To this end, this work proposes a victim-side pseudo-reference supervision mechanism that operates without benchmark feedback. Specifically, it leverages clean continuations of unlabeled instructions as pseudo-references, integrating local search with preference optimization to train a utility-degrading prefix generator. This research pioneers a pure utility attack paradigm devoid of harmful content, achieving generative prefix learning through white-box search and reward refinement. Extensive evaluations across four models and seven benchmarks demonstrate an average utility degradation of 26.8 percentage points, with significant performance deterioration observed in 27 out of 28 metrics, effectively revealing critical utility weaknesses in large language models.
📝 Abstract
Prompt injection can degrade benign task performance without eliciting harmful content. Yet many attack objectives depend on task labels or predefined target responses. We present ENDOPROMPT, a white-box method that learns utility-degrading prefixes from unlabeled instructions. Its generator takes the request text as input. Clean victim continuations serve as pseudo-references: local search identifies prefixes that reduce continuation likelihood, and preference fitting on comparisons within the same instruction, followed by reward refinement, distills this signal into a generator. At deployment, the generator produces one prefix per request without further victim-side search. Across four instruction-tuned models and the complete splits of seven benign benchmarks, ENDOPROMPT yields a mean utility change of -26.8 percentage points; 27 of 28 cells are negative. Failure analysis reveals output expansion and prefix reuse; the controls do not establish a degradation advantage from request matching. Victim-derived supervision can reveal utility weaknesses without benchmark feedback or prescribed failure responses. The code will be released upon acceptance.