Researchers test how language models handle numerical formats in word problems
A new paper on arXiv examines whether language models consistently answer numerical word problems regardless of how quantities are expressed. The authors created 3,600 exact-rational problems and 8,600 prompts across five transformation types, then tested five open-weight models. After normalizing answers, the models scored between 0.969 and 0.996 on canonical accuracy but dropped to 0.848–0.981…
Key points
- Researchers generated 3,600 exact-rational and 8,600 prompts testing numerical format invariance in language models
- Five open-weight models scored 0.969–0.996 on canonical accuracy but dropped to 0.848–0.981 on orbit correctness
- Mistral Small 4 scored 0.699 on unit-converted inputs, with 265 errors differing by exact powers of ten
The study also found that representation consensus did not outperform paraphrase consensus in a 9,000-call experiment. The paper includes a benchmark, evaluation records, and raw responses, all available in an ancillary archive.
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- Study finds wiki-indexed LLM outperforms vector RAG for cross-unit questions in ML courses · 1 src
- Study proposes skill habits to fix AI agent inconsistency · 1 src
- Researchers propose latent equivalence learning for enterprise data agents, score 94.67% on benchmark · 1 src
- Researchers extract circuits from language models using Attention routing · 1 src
- ReAdapt improves warm‑introduction and reaction selection accuracy for Gemini‑3‑Flash · 1 src
Comments
via GitHub Discussions