Simon Willison tests Qwen3.8-27B arithmetic in words
Simon Willison replicated a two-year-old experiment by Colin Frasier, which tested GPT-4o's ability to sum large numbers and return the result in words. Willison ran the test on local hardware, specifically a DGX Spark, using the Qwen3.8-27B-Q4KM.gguf model. He utilized a Codex Remote session with GPT-6 Astra to orchestrate the experiment, pasting Frasier's original chart into the system to…
Key points
- Willison tested Qwen3.8-27B-Q4KM.gguf on a DGX Spark to sum large numbers in words.
- The model achieved 167 out of 169 correct answers in a one-shot run with reasoning enabled.
- GPT-6 Astra was used via Codex Remote to orchestrate the experiment on local hardware.
The experiment involved 30 attempts per number combination with reasoning disabled, followed by a second run with reasoning enabled. In the reasoning-enabled phase, Willison ran only one sample per pair due to the increased time required. The model correctly calculated the sum in words for 167 out of 169 attempts. Willison noted that because these were one-shot runs, a second iteration would likely yield different results. The report includes reasoning traces showing the model aligning digits and adding from right to left, occasionally correcting its own alignment errors.
Model page: GPT-6 Astra →
Qwen3.8 27B addition in words
Simon Willison · 4 October 2026
Loading the full article…
This text was published by Simon Willison and written by Simon Willison. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- Researchers present ProFormer for classifying proteomic data · 1 src
- Researchers track Chinese AI agent fleet that may run on Tencent, targeting Alibaba's Amap · 1 src
- Deep learning pipeline differentiates HCM from cardiac amyloidosis · 1 src
- ResearchAndMarkets.com releases Argentina data center market report projecting value to USD 825 million · 1 src
- Import AI 475 covers swarm scaling, DeepMind biology watermarks, and AI science economy · 1 src
Comments
via GitHub Discussions