DigestAI news desk

Language Models Can't Detect Their Own Training Data

A new study published on arXiv cs.CL reveals that language models struggle to detect sentences in their training data even when they are unusually easy to predict. Two model families, OLMo-2 and Pythia, have released their pretraining corpora for public scrutiny. By comparing the frequency of sentences appearing across these corpora, researchers can determine whether a sentence was part of the…

1 source primary source

Key points

  • Two model families released their pretraining corpora
  • Language models struggle to detect sentences from their own training data
  • Models with up to 13 billion parameters show a faint trace of exposure
Read the original at arXiv cs.CL · by Arman Nik Khah primary source Open source ↗
Topics · follow one to build your own front page
OLMo-2Pythia

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.

Comments

via GitHub Discussions

More in Generative AI & Models

All →

Related stories