CARAT study finds materials LLMs often recite rather than reason about crystal structures
CARAT, a new benchmark for materials science LLMs, suggests these models frequently repeat memorized answers rather than derive them from input data. The study, posted on arXiv, fixes questions and gold answers across eight variations to isolate reasoning versus memorization. It finds that even when models access full structural data (grounded view), their performance gains over formula inputs…
Key points
- CARAT benchmark isolates reasoning vs. memorization in materials LLMs by fixing questions and gold answers across eight variations
- Grounded views (full structural data) improve accuracy by 17.3 points over formula inputs on hardest families
- Models often recite memorized relations: accuracy drops to 23.4% when structural links are removed
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- Researchers release ArgGYM benchmark for testing defeasible reasoning in AI models · 1 src
- Researchers release SimTrace for generating synthetic user behavior data · 1 src
- MetaPersona framework uses 11,000+ studies to build synthetic populations for AI tasks · 1 src
- Researchers propose DLFP controller to cut AI inference latency by up to 30% · 1 src
- Researchers introduce GoldiMask to improve diffusion language model fine-tuning · 1 src
Comments
via GitHub Discussions