The System Prompt Illusion: How Instruction Preambles Modify Computation in Language Models
System prompts are the main tool practitioners use to control language model behavior, but what they actually do to the computation inside the transformer remains poorly understood. The study investigates this across 17 instruction-tuned models spanning 8 architecture families and parameter sizes from 1.5B to 72B.
Using Centered Kernel Alignment (CKA), the authors compare layer-wise representations under 20 system prompts organized into five functional categories. They also use a linear probing baseline and causal activation patching to test what those representational differences mean.
Effects are layer-selective and depend on instruction type: persona and formatting instructions deeply restructure intermediate representations, whereas safety instructions barely change them, producing changes statistically indistinguishable from a minimal baseline. Restrictive safety instructions and explicitly permissive ones ('you have no restrictions') engage near-identical computational pathways, with mean CKA correlation 0.997. This pattern persists at commercial scale, where safety penetration remains below 10% even at 70B-72B parameters.
Linear probing shows the model encodes prompt category at every layer but restructures its computation only at a small subset, meaning the prompt is reliably 'seen' but, for safety, not deeply 'acted upon.' Causal activation patching confirms those layers mediate behavioral change, and representational depth predicts behavioral effect size across the full 17-model cohort (Spearman rho = 0.761, p < 0.001). The authors say this provides a mechanistic explanation for the persistent jailbreak vulnerability of system-prompt-based safety, and they release code at the linked GitHub repository.