{"version":1,"type":"story","url":"https://digestai.news/story/researchers-test-how-language-models-handle-numerical-formats-in-word","json":"https://digestai.news/story/researchers-test-how-language-models-handle-numerical-formats-in-word.json","markdown":"https://digestai.news/story/researchers-test-how-language-models-handle-numerical-formats-in-word.md","slug":"researchers-test-how-language-models-handle-numerical-formats-in-word","headline":"Researchers test how language models handle numerical formats in word problems","summary":"A new paper on arXiv examines whether language models consistently answer numerical word problems regardless of how quantities are expressed. The authors created 3,600 exact-rational problems and 8,600 prompts across five transformation types, then tested five open-weight models. After normalizing answers, the models scored between 0.969 and 0.996 on canonical accuracy but dropped to 0.848–0.981 on orbit correctness and invariance. Mistral Small 4 struggled with unit-converted inputs, scoring 0.699 and producing 265 errors off by exact powers of ten.\n\nThe study also found that representation consensus did not outperform paraphrase consensus in a 9,000-call experiment. The paper includes a benchmark, evaluation records, and raw responses, all available in an ancillary archive.","keyPoints":["Researchers generated 3,600 exact-rational and 8,600 prompts testing numerical format invariance in language models","Five open-weight models scored 0.969–0.996 on canonical accuracy but dropped to 0.848–0.981 on orbit correctness","Mistral Small 4 scored 0.699 on unit-converted inputs, with 265 errors differing by exact powers of ten"],"whyItMatters":"The findings highlight persistent flaws in how models handle numerical reasoning, especially unit conversions, which could affect real-world applications like finance or science where precision matters.","category":{"slug":"research","name":"Research","url":"https://digestai.news/category/research"},"entities":{"companies":[],"models":["Mistral Small 4"],"people":[]},"firstPublishedAt":"2026-09-23T04:00:00Z","updatedAt":"2026-09-23T04:00:00Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"arXiv cs.CL","title":"Same Quantity, Different Answer: Numerical Representation Invariance in Language Models","url":"https://arxiv.org/abs/2609.25009","publishedAt":"2026-09-23T04:00:00Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[],"thread":null,"cite":{"text":"Digest AI, \"Researchers test how language models handle numerical formats in word problems\", 23 September 2026, https://digestai.news/story/researchers-test-how-language-models-handle-numerical-formats-in-word","publisher":"Digest AI","title":"Researchers test how language models handle numerical formats in word problems","datePublished":"2026-09-23T04:00:00Z","url":"https://digestai.news/story/researchers-test-how-language-models-handle-numerical-formats-in-word"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}