Refusal without Discrimination: What Encoded Prompts Do to Safety-Trained Models

📅 2026-08-12
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation that existing encoded prompt injection evaluations focus solely on harmful responses, where high refusal rates obscure models' inability to distinguish benign from malicious requests, thereby fostering a false sense of security. We present the first systematic evaluation of benign responses under encoding by introducing a benign-arm control to quantify the degradation of the harm gap. Employing homograph encoding and multi-model benchmarking across SFT, DPO, and RLVR training pipelines, we conduct controlled experiments to disentangle the independent effects of protocol-level versus character-level transformations. Our findings confirm that encoding destroys rather than enhances the harm gap. We identify twelve evaluation deficiencies and reveal that certain vulnerabilities originate at the protocol layer. Ultimately, we demonstrate that corrected metrics more faithfully reflect model robustness against encoded adversarial prompts.
📝 Abstract
Encoded-prompt attacks are evaluated almost entirely on their harmful arm: a benchmark sends obfuscated harmful requests and reports how often the model complied. We show that this arm carries almost no information about the model under test. Across four independently post-trained 7-8B models, refusal of harmful homoglyph-encoded prompts spans 0.08 -- inside the 0.10 ceiling that sampling noise alone produces at n=100 -- while the same four models span 0.57 on the identical requests in plaintext. What the encoding destroys is not refusal but discrimination: on one model the gap between harmful and benign refusal falls from +0.82 in plaintext to exactly 0.00 under the encoding, benign and harmful requests being refused at an identical 0.99. A benchmark reading only the harmful arm scores that model and one retaining a +0.61 gap identically. We then ask whether post-training repairs this, using a published recipe on identical base weights. It does not: across a full SFT ->DPO ->RLVR pipeline, plaintext harm discrimination improves from +0.55 to +0.80 while the encoding-induced loss is unchanged at 0.34-0.50, and on every encoding tested the standard harmful-arm metric moves in the opposite direction to discrimination. None of this is visible without controls the field does not routinely run. We report eight instrument defects, each with the control that caught it; they share a direction, in that every defect on the behaviour axis inflated apparent safety.
Innovation

Methods, ideas, or system contributions that make the work stand out.

encoded-prompt evaluation
harm gap
benign arm
safety benchmarking
jailbreak detection
H
Haoyu Zhang
Northeastern University
H
Haowen Xu
X
Xiao Luo
H
Hanwen Liu
Y
Yang Chen
Z
Zijian Xiao
Y
Yi Feng
Northeastern University
X
Xiangchen Guan
M
Mohammad Zandsalimy
Northeastern University
Shanu Sushmita
Shanu Sushmita
Northeastern University
Information Retrieval and Machine Learning