# Study shows system prompts barely change computation in safety instructions across 17 language models

Digest AI · Research · published 2026-10-01T04:00:00Z

Canonical: https://digestai.news/story/study-shows-system-prompts-barely-change-computation-in-safety-instruc

## Summary

Researchers analyzed 17 instruction-tuned language models from 8 architecture families, ranging from 1.5B to 72B parameters, to understand how system prompts affect internal computation. Using Centered Kernel Alignment (CKA), they found that persona and formatting prompts significantly restructure layer-wise representations, while safety instructions produce changes statistically indistinguishable from a minimal baseline. Restrictive and permissive safety prompts engage nearly identical computational pathways, with mean CKA correlation of 0.997, even at commercial scale where safety penetration remains below 10% in 70B-72B models. The study indicates that models encode prompt category at every layer but only alter computation in a small subset of layers, explaining why system-prompt-based safety is vulnerable to jailbreaks. Code for the study is available on GitHub.

## Key points

- Safety instructions produce representation changes statistically indistinguishable from baseline across 17 models
- Restrictive and permissive safety prompts show near-identical computational pathways (CKA correlation 0.997)
- Safety penetration remains below 10% even in 70B-72B parameter models

## Why it matters

The study reveals a core limitation in current AI safety practices: system prompts do not deeply alter model computation for safety, explaining why jailbreaks persist and suggesting a need for stronger, intrinsic safety mechanisms.

## Sources

1. [The System Prompt Illusion: How Instruction Preambles Modify Computation in Language Models](https://arxiv.org/abs/2609.38205) (arXiv cs.CL, 2026-10-01, primary source)

## Cite

Digest AI, "Study shows system prompts barely change computation in safety instructions across 17 language models", 1 October 2026, https://digestai.news/story/study-shows-system-prompts-barely-change-computation-in-safety-instruc

---

Written by Digest AI's editorial model from the linked sources; the sources are the record. Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse
JSON: https://digestai.news/story/study-shows-system-prompts-barely-change-computation-in-safety-instruc.json
