Decision Hijacking: Prompt Injection Attacks on Jev's Typed Probabilistic Decisions

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the underexplored impact of prompt injection on non-generative decision models subject to schema constraints. Focusing on the Jev decision model, we reconstruct a test set based on InjecAgent cases and systematically evaluate how malicious content perturbs action probabilities through adaptive attack optimization combined with statistical validation. Our findings reveal that while schema constraints restrict the target selection space, they fail to eliminate injection risks, necessitating particular attention to decision shifts within the permissible action set. Experimental results demonstrate that adaptive attacks double the mean probability of target actions, increasing the success rate of novel validation invocations from 1.8% to 3.5%. These outcomes underscore significant security vulnerabilities inherent in schema-constrained decision models when exposed to adversarial prompt injections.
📝 Abstract
Most studies of prompt injection focus on generative agents, leaving their effects on models with schema-defined outputs unclear. We examine these effects in Jev, a non-generative decision model, using 510 reconstructed InjecAgent cases. Malicious content shifts action probabilities but rarely causes Jev to select the attacker's target. Override markers reduce this influence, while claims of contextual relatedness have small effects. Adaptive attacks using score feedback double the mean highest attacker-target probability found during optimization, while success on fresh validation calls rises from 1.8% to 3.5%. Exploratory analysis links these successes to small initial decision margins or greater attacker control over the observation. Together, these findings show that schema-defined outputs change but do not eliminate prompt-injection risk, highlighting the need to evaluate how untrusted content influences choices within the allowed action set.
Problem

Research questions and friction points this paper is trying to address.

Prompt Injection
Decision Hijacking
Schema-defined Outputs
Probabilistic Decisions
Adversarial Attacks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Prompt Injection
Decision Hijacking
Schema-defined Outputs
Adaptive Attacks
Typed Probabilistic Decisions
🔎 Similar Papers
No similar papers found.