{"version":1,"type":"story","url":"https://digestai.news/story/simon-willison-tests-qwen3-8-27b-arithmetic-in-words","json":"https://digestai.news/story/simon-willison-tests-qwen3-8-27b-arithmetic-in-words.json","markdown":"https://digestai.news/story/simon-willison-tests-qwen3-8-27b-arithmetic-in-words.md","slug":"simon-willison-tests-qwen3-8-27b-arithmetic-in-words","headline":"Simon Willison tests Qwen3.8-27B arithmetic in words","summary":"Simon Willison replicated a two-year-old experiment by Colin Frasier, which tested GPT-4o's ability to sum large numbers and return the result in words. Willison ran the test on local hardware, specifically a DGX Spark, using the Qwen3.8-27B-Q4KM.gguf model. He utilized a Codex Remote session with GPT-6 Astra to orchestrate the experiment, pasting Frasier's original chart into the system to guide the process.\n\nThe experiment involved 30 attempts per number combination with reasoning disabled, followed by a second run with reasoning enabled. In the reasoning-enabled phase, Willison ran only one sample per pair due to the increased time required. The model correctly calculated the sum in words for 167 out of 169 attempts. Willison noted that because these were one-shot runs, a second iteration would likely yield different results. The report includes reasoning traces showing the model aligning digits and adding from right to left, occasionally correcting its own alignment errors.","keyPoints":["Willison tested Qwen3.8-27B-Q4KM.gguf on a DGX Spark to sum large numbers in words.","The model achieved 167 out of 169 correct answers in a one-shot run with reasoning enabled.","GPT-6 Astra was used via Codex Remote to orchestrate the experiment on local hardware."],"whyItMatters":"This test highlights the specific limitations of local, quantized models in precise arithmetic tasks, even when reasoning capabilities are enabled. It provides a concrete benchmark for the reliability of open-weight models in controlled, low-resource environments.","category":{"slug":"research","name":"Research","url":"https://digestai.news/category/research"},"entities":{"companies":["OpenAI"],"models":["Qwen3.8-27B-Q4KM.gguf","GPT-4o","GPT-6 Astra"],"people":["Simon Willison","Colin Frasier"]},"firstPublishedAt":"2026-10-04T23:34:00Z","updatedAt":"2026-10-04T23:34:00Z","sourceCount":1,"hasPrimarySource":false,"sources":[{"outlet":"Simon Willison","title":"Qwen3.8 27B addition in words","url":"https://simonwillison.net/2026/Oct/4/qwen38-addition-in-words","publishedAt":"2026-10-04T23:34:00Z","type":"newsletter","primary":false,"lead":true}],"sourceNotes":null,"discussions":[],"thread":null,"cite":{"text":"Digest AI, \"Simon Willison tests Qwen3.8-27B arithmetic in words\", 4 October 2026, https://digestai.news/story/simon-willison-tests-qwen3-8-27b-arithmetic-in-words","publisher":"Digest AI","title":"Simon Willison tests Qwen3.8-27B arithmetic in words","datePublished":"2026-10-04T23:34:00Z","url":"https://digestai.news/story/simon-willison-tests-qwen3-8-27b-arithmetic-in-words"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}