DigestAI news desk

Cut through the AI noise.

Research

Researchers test how language models handle numerical formats in word problems

A new paper on arXiv examines whether language models consistently answer numerical word problems regardless of how quantities are expressed. The authors created 3,600 exact-rational problems and 8,600 prompts across five transformation types, then tested five open-weight models. After normalizing answers, the models scored between 0.969 and 0.996 on canonical accuracy but dropped to 0.848–0.981…

1 source primary source

Key points

  • Researchers generated 3,600 exact-rational and 8,600 prompts testing numerical format invariance in language models
  • Five open-weight models scored 0.969–0.996 on canonical accuracy but dropped to 0.848–0.981 on orbit correctness
  • Mistral Small 4 scored 0.699 on unit-converted inputs, with 265 errors differing by exact powers of ten

The study also found that representation consensus did not outperform paraphrase consensus in a 9,000-call experiment. The paper includes a benchmark, evaluation records, and raw responses, all available in an ancillary archive.

Read the original at arXiv cs.CL · by Ephraim Atta-Duncan primary sourceOpen source ↗
Topics · follow one to build your own front page
Mistral Small 4

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Research

All →

Related stories