Study shows system prompts barely change computation in safety instructions across 17 language models
Researchers analyzed 17 instruction-tuned language models from 8 architecture families, ranging from 1.5B to 72B parameters, to understand how system prompts affect internal computation. Using Centered Kernel Alignment (CKA), they found that persona and formatting prompts significantly restructure layer-wise representations, while safety instructions produce changes statistically…
Key points
- Safety instructions produce representation changes statistically indistinguishable from baseline across 17 models
- Restrictive and permissive safety prompts show near-identical computational pathways (CKA correlation 0.997)
- Safety penetration remains below 10% even in 70B-72B parameter models
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- A Holistic Assessment of the Carbon Footprint of Noor, a Very Large Arabic Language Model · 1 src
- Korean legal study finds KLUE-BERT outperforms GPT models in sexual offense text classification · 1 src
- Researchers question human-derived bias measures for LLM evaluation · 1 src
- arXiv study finds reading LLM judges from first token overstates position bias · 1 src
- Study compares On-Device NER models for speed, cost and accuracy · 1 src
Comments
via GitHub Discussions