DigestAI news desk

Cut through the AI noise.

Research

Study identifies under 1% of BERT neurons driving AI-Text detection

Researchers examined a frozen BERT‑base‑uncased encoder to determine which neurons support AI‑generated text detection. Using the RAID benchmark across six generators, they applied an L1‑to‑L2 sparse‑probing protocol to all 9,216 CLS hidden‑state dimensions. The study found a stable set of under 1% neurons per generator that largely preserves detection accuracy.

1 source primary source

Key points

  • under 1% of bert neurons per generator support ai‑text detection
  • activation patching flips predictions an order of magnitude more often than random sets
  • selected neurons retain 86‑94% accuracy on unseen generator families

Activation patching confirmed the causal relevance of this set, flipping predictions an order of magnitude more often than size‑matched random sets. Mean‑ablating the same neurons left accuracy largely intact, indicating redundancy. Cross‑generator analysis revealed a bipartite structure: instruction‑tuned generators concentrate 30‑36% of stable neurons in the final layer, whereas base generators fall below 14%.

Leave‑one‑family‑out evaluation showed the selected neurons retain 86‑94% of the full‑feature ceiling on unseen generator families, suggesting a detector can operate on a small fixed subspace without re‑identifying neurons per generator.

Read the original at arXiv cs.CL · by Pawe{\l} Blicharz, Mi{\l}osz Grunwald primary sourceOpen source ↗
Topics · follow one to build your own front page
bert‑base‑uncasedgurnee et al. (2023)

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Research

All →

Related stories