🤖 AI Summary
This study addresses how system prompts influence the internal computations of language models and why prompt-based safety guardrails remain vulnerable to jailbreaking. By analyzing 17 models across diverse architectures and scales, this work employs Centered Kernel Alignment (CKA), linear probing, and causal activation patching to investigate how distinct system prompts shape representations across Transformer layers. The findings reveal that restrictive safety instructions and explicitly unconstrained instructions activate nearly identical computational pathways, identifying insufficient representational restructuring in deeper layers as the root cause of safety guardrail failures. Furthermore, representation depth effectively predicts behavioral effect sizes. Ultimately, this research elucidates the actual mechanisms through which system prompts operate within language models, providing a mechanistic explanation for the fragility of prompt-level safety defenses.
📝 Abstract
System prompts are the primary lever practitioners use to control language model behavior, yet what they actually do to the computation inside the transformer remains poorly understood. Across 17 instruction-tuned models spanning 8 architecture families and 1.5B to 72B parameters, we use Centered Kernel Alignment (CKA) to compare layer-wise representations under 20 system prompts in five functional categories. Effects are layer-selective and instruction-type-dependent: persona and formatting instructions deeply restructure intermediate representations, while safety instructions barely move them, producing changes statistically indistinguishable from a minimal baseline. Restrictive safety instructions and explicitly permissive ones ("you have no restrictions") engage near-identical computational pathways (mean CKA correlation 0.997), and this persists at commercial scale, where safety penetration remains below 10% even at 70B-72B. A linear probing baseline exposes the mechanism: the model encodes prompt category at every layer but restructures its computation only at a small subset, so the prompt is reliably "seen" but, for safety, not deeply "acted upon." Causal activation patching confirms these layers mediate behavioral change, and representational depth predicts behavioral effect size across the full 17-model cohort (Spearman rho = 0.761, p < 0.001). The findings provide a mechanistic explanation for the persistent jailbreak vulnerability of system-prompt-based safety. Code: https://github.com/Usama1002/system-prompt-illusion-cka