{"version":1,"type":"story","url":"https://digestai.news/story/study-identifies-under-1-of-bert-neurons-driving-ai-text-detection","json":"https://digestai.news/story/study-identifies-under-1-of-bert-neurons-driving-ai-text-detection.json","markdown":"https://digestai.news/story/study-identifies-under-1-of-bert-neurons-driving-ai-text-detection.md","slug":"study-identifies-under-1-of-bert-neurons-driving-ai-text-detection","headline":"Study identifies under 1% of BERT neurons driving AI-Text detection","summary":"Researchers examined a frozen BERT‑base‑uncased encoder to determine which neurons support AI‑generated text detection. Using the RAID benchmark across six generators, they applied an L1‑to‑L2 sparse‑probing protocol to all 9,216 CLS hidden‑state dimensions. The study found a stable set of under 1% neurons per generator that largely preserves detection accuracy.\n\nActivation patching confirmed the causal relevance of this set, flipping predictions an order of magnitude more often than size‑matched random sets. Mean‑ablating the same neurons left accuracy largely intact, indicating redundancy. Cross‑generator analysis revealed a bipartite structure: instruction‑tuned generators concentrate 30‑36% of stable neurons in the final layer, whereas base generators fall below 14%.\n\nLeave‑one‑family‑out evaluation showed the selected neurons retain 86‑94% of the full‑feature ceiling on unseen generator families, suggesting a detector can operate on a small fixed subspace without re‑identifying neurons per generator.","keyPoints":["under 1% of bert neurons per generator support ai‑text detection","activation patching flips predictions an order of magnitude more often than random sets","selected neurons retain 86‑94% accuracy on unseen generator families"],"whyItMatters":"Understanding which neurons drive ai‑text detection can reduce model size, improve robustness, and guide future detector design.","category":{"slug":"research","name":"Research","url":"https://digestai.news/category/research"},"entities":{"companies":[],"models":["bert‑base‑uncased"],"people":["gurnee et al. (2023)"]},"firstPublishedAt":"2026-09-28T04:00:00Z","updatedAt":"2026-09-28T04:00:00Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"arXiv cs.CL","title":"A Mechanistic Study of AI-Text Detection Neurons in Frozen BERT: Sparse Probing and Activation Patching on RAID","url":"https://arxiv.org/abs/2609.30287","publishedAt":"2026-09-28T04:00:00Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[],"thread":null,"cite":{"text":"Digest AI, \"Study identifies under 1% of BERT neurons driving AI-Text detection\", 28 September 2026, https://digestai.news/story/study-identifies-under-1-of-bert-neurons-driving-ai-text-detection","publisher":"Digest AI","title":"Study identifies under 1% of BERT neurons driving AI-Text detection","datePublished":"2026-09-28T04:00:00Z","url":"https://digestai.news/story/study-identifies-under-1-of-bert-neurons-driving-ai-text-detection"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}