DigestAI news desk

Cut through the AI noise.

Research

CARAT study finds materials LLMs often recite rather than reason about crystal structures

CARAT, a new benchmark for materials science LLMs, suggests these models frequently repeat memorized answers rather than derive them from input data. The study, posted on arXiv, fixes questions and gold answers across eight variations to isolate reasoning versus memorization. It finds that even when models access full structural data (grounded view), their performance gains over formula inputs…

1 source primary source

Key points

  • CARAT benchmark isolates reasoning vs. memorization in materials LLMs by fixing questions and gold answers across eight variations
  • Grounded views (full structural data) improve accuracy by 17.3 points over formula inputs on hardest families
  • Models often recite memorized relations: accuracy drops to 23.4% when structural links are removed
Read the original at arXiv cs.AI · by Jiajun Wu, Jian Yang, Zixiang Ni, Zhenzhu Li, Bin Chong primary sourceOpen source ↗
Topics · follow one to build your own front page
CARAT

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Research

All →

Related stories