{"version":1,"type":"story","url":"https://digestai.news/story/swift-1-5-qwen3-8-27b-oq8e-mtp-achieves-34-8-tokens-per-second-on-appl","json":"https://digestai.news/story/swift-1-5-qwen3-8-27b-oq8e-mtp-achieves-34-8-tokens-per-second-on-appl.json","markdown":"https://digestai.news/story/swift-1-5-qwen3-8-27b-oq8e-mtp-achieves-34-8-tokens-per-second-on-appl.md","slug":"swift-1-5-qwen3-8-27b-oq8e-mtp-achieves-34-8-tokens-per-second-on-appl","headline":"Swift-1.5-Qwen3.8-27b-oQ8e-mtp achieves 34.8 tokens per second on Apple M5 Max","summary":"A Reddit user shared benchmark results for the **Swift-1.5-Qwen3.8-27b-oQ8e-mtp** model on an **Apple M5 Max** device on September 29, 2026. The model reached **34.8 tokens per second** in a coding scenario, generating a playable game in a single HTML file without edits. The benchmark tool and parameters are detailed for reproducibility, though no official confirmation or context from **Qwen** or **Apple** is provided.","keyPoints":["Swift-1.5-Qwen3.8-27b-oQ8e-mtp hit **34.8 tokens per second** on Apple M5 Max in a coding task","Median speed across five runs was **34.5 tokens per second**, with **KV cache** settings affecting results"],"whyItMatters":"Faster inference on Apple silicon could lower costs for developers running large models locally, but benchmarks from community sources lack official validation.","category":{"slug":"research","name":"Research","url":"https://digestai.news/category/research"},"entities":{"companies":["Apple","Qwen"],"models":["Swift-1.5-Qwen3.8-27b-oQ8e-mtp"],"people":[]},"firstPublishedAt":"2026-09-29T16:07:22Z","updatedAt":"2026-09-29T16:07:22Z","sourceCount":1,"hasPrimarySource":false,"sources":[{"outlet":"llm-bench.io","title":"Swift-1.5-Qwen3.8-27b-oQ8e-mtp on Apple M5 Max — 34.8 tok/s — llm-bench.io","url":"https://llm-bench.io/benchmarks/cmumgubnk004o01o00kqgaqcv","publishedAt":"2026-09-29T16:07:22Z","type":"press","primary":false,"lead":true}],"sourceNotes":null,"discussions":[{"site":"Reddit","url":"https://www.reddit.com/r/LocalLLaMA/comments/1wte7n0/swift15qwen3827boq8emtp_on_apple_m5_max_348_toks/","points":null}],"thread":{"title":"Local AI Inference Speeds Hit New Peaks","url":"https://digestai.news/thread/llama-cpp-achieves-up-to-42x-faster-prompt-lookup-drafting","storyCount":2},"cite":{"text":"Digest AI, \"Swift-1.5-Qwen3.8-27b-oQ8e-mtp achieves 34.8 tokens per second on Apple M5 Max\", 29 September 2026, https://digestai.news/story/swift-1-5-qwen3-8-27b-oq8e-mtp-achieves-34-8-tokens-per-second-on-appl","publisher":"Digest AI","title":"Swift-1.5-Qwen3.8-27b-oQ8e-mtp achieves 34.8 tokens per second on Apple M5 Max","datePublished":"2026-09-29T16:07:22Z","url":"https://digestai.news/story/swift-1-5-qwen3-8-27b-oq8e-mtp-achieves-34-8-tokens-per-second-on-appl"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}