Study finds brain-alignment and Cross-Lingual scores can mislead when probes fail
A new arXiv paper challenges the use of similarity scores as evidence that language models share structure with the human brain or across languages. The authors examine two common settings. In cross-lingual transfer, a probe trained to distinguish grammatical from ungrammatical sentences in one language transfers worse to more distant languages, but the probe itself degrades along the same axis.…
Key points
- Brain-alignment score rises from 0.10 to 0.34 but destroyed-target model still scores 0.31
- Statistical significance flips from p=0.0006 to p=0.155 when using languages not pairs as units
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- Anthropic study: task understanding beats job title for AI success · 1 src
- Researchers test fixed token codes for language models at 100B-token scale · 1 src
- Researchers propose surrogate log-probabilities to audit LLM agents · 3 src
- Behavioral history outperforms descriptions for LLM synthetic personas · 1 src
- Researchers introduce ROAR to unify AI-driven research system runs · 1 src
Comments
via GitHub Discussions