Unread or Unenforced? Separating Representation from Enforcement Failure in Content Guards

๐Ÿ“… 2026-08-12
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study investigates whether content moderation failures stem from attacks remaining unrepresented or from represented attacks evading interception. To disentangle representation from execution, we propose a โ€œhonest readoutโ€ framework employing residual stream probes. We rigorously exclude false positives through length-matched null hypotheses, permutation tests, and plaintext-versus-encoded control experiments. Empirically, our analysis reveals strategic safety failures under specific encodings while finding no evidence of purely format-based detectors. By precisely delineating the failure boundaries and applicability scope of existing defense mechanisms, this work provides a principled diagnostic methodology for understanding how encoding-based jailbreaks circumvent alignment safeguards in large language models.
๐Ÿ“ Abstract
When an encoded attack passes a content guard, the guard either never represented the payload's harmful content or represented it and failed to act. End-to-end attack success rate reports one number for both, yet the remedies are opposite: one is a representational limit that more safety training cannot reach, the other a decision rule that it can. We separate them by reading a guard's own residual stream -- a content probe fitted on plaintext and transferred, without refitting, to the encoded condition -- alongside its verdict logits, at the cost of one forward pass and no judge model. Doing this honestly is most of the problem, and it is our main contribution. A permutation test licenses the decode measurement on 17 of 19 conditions for one open guard and 12 of 19 for another; a length-matched null and a control floor calibrated on conditions the guard's base model provably cannot decode reduce both to 4. The discarded cells are not marginal ones: the largest result in our first analysis -- a guard representing a cipher at AUROC 0.72 while blocking none of it -- is an artefact on an encoding its base model decodes at rate zero. On one guard, two screens sharing no input agree exactly on which conditions to reject. What survives is a policy failure that is real but narrower than the uncontrolled analysis claimed: 7 to 23 per 100 prompts represented and not blocked on conditions the guards block heavily, and 56 per 100 on one condition a guard barely blocks. Blocked without decoding is near zero throughout, so neither guard reacts to the appearance of encoding rather than to content. On genuine ciphers both guards block essentially nothing, and we report those cells as unmeasured rather than as evidence of failure to decode.
Problem

Research questions and friction points this paper is trying to address.

content guard
encoded attack
representation failure
enforcement failure
AI safety
Innovation

Methods, ideas, or system contributions that make the work stand out.

residual stream probing
failure disentanglement
permutation test
content guards
encoded attacks
๐Ÿ”Ž Similar Papers
No similar papers found.
๐Ÿ’ผ Related Jobs
No related jobs found.
H
Haoyu Zhang
Northeastern University
Y
Yi Feng
Northeastern University
S
Shibo Zheng
H
Hanwen Liu
H
Haowen Xu
X
Xiao Luo
Z
Zhuoxi Wang
M
Mohammad Zandsalimy
Northeastern University
Shanu Sushmita
Shanu Sushmita
Northeastern University
Information Retrieval and Machine Learning