AI fails to rediscover relativity in 'Einstein test' experiments
Researchers are testing whether large language models can replicate scientific breakthroughs by training them on historical data cut off before major discoveries. This concept, dubbed the "Einstein test" by Google DeepMind co-founder Demis Hassabis, aims to determine if AI can perform the abductive reasoning required for paradigm shifts, such as deriving general relativity from pre-1911 knowledge.
Key points
- Google DeepMind proposed the 'Einstein test' to check if AI can rediscover general relativity using pre-1911 data.
- Experiments show current LLMs fail to derive fundamental physics laws, often generating incorrect or probabilistic theories.
- Researchers argue AI lacks abductive reasoning, making it unsuitable for paradigm-shifting scientific breakthroughs without new architectures.
Early attempts have revealed significant limitations. A model trained on pre-1900 data by independent researcher Michael Hla showed only "glimpses of intuition" regarding quantum mechanics but failed to grasp underlying physics, often producing plausible-sounding but incorrect statements. Similarly, a foundation model tested by MIT researchers on synthetic orbital data failed to infer Newton's law of gravitation, instead generating unique, incorrect laws for each planetary system. These results suggest current AI excels at pattern matching and grunt work but lacks the creative leap necessary for transformative scientific discovery.
Other teams, including those at the University of Zurich, are building "vintage" models with specific historical cut-offs to test for "sparks of genius." While data leakage remains a technical hurdle, experts argue that while AI may eventually make non-trivial discoveries, it currently cannot distinguish correct theories from probabilistic noise without external verification methods.
Coverage and discussion
1 source- Hacker News discussion · 10 points news.ycombinator.com
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.
Comments
via GitHub Discussions