{"version":1,"type":"story","url":"https://digestai.news/story/gufo-qwen3-6-35b-a3b-q6dense-3095tok-s-prefill-190-tok-s-decode-on-str","json":"https://digestai.news/story/gufo-qwen3-6-35b-a3b-q6dense-3095tok-s-prefill-190-tok-s-decode-on-str.json","markdown":"https://digestai.news/story/gufo-qwen3-6-35b-a3b-q6dense-3095tok-s-prefill-190-tok-s-decode-on-str.md","slug":"gufo-qwen3-6-35b-a3b-q6dense-3095tok-s-prefill-190-tok-s-decode-on-str","headline":"gufo-Qwen3.6-35B-A3B-Q6dense - 3095tok/s prefill; 190 tok/s decode on Strix Halo","summary":"The Gufo project posted benchmark results for its vertical local inference engine built for AMD Strix Halo hardware (Ryzen AI MAX+ 395 with Radeon 8060S, up to 128 GiB unified memory). The engine uses speculative decoding with a DFlash2 draft model and targets rootless Podman containers on Linux x86-64 with ROCm 7.2.3. Gufo provides an OpenAI-compatible API at localhost:8080/v1 and supports text, images, tools, and streaming. The project emphasizes quality-preserving optimization, per-model kernels, and concurrent request handling. Windows and dual-Strix RDMA variants exist but are not yet merged. Model weights are downloaded separately from Hugging Face (unsloth and z-lab repos). Build requires GCC 15.3, CMake 3.21+, Ninja, and ROCm development packages. Gufo's code is MIT licensed; model weights retain their publishers' terms.","keyPoints":["Runs Qwen3.6-35B-A3B-Q6dense with DFlash2 speculative decoding via Podman","OpenAI-compatible API on localhost:8080/v1; MIT-licensed code, separate model weights"],"whyItMatters":"Shows high-throughput local LLM inference on consumer AMD APUs with unified memory, enabling private, offline AI workloads without discrete GPUs.","category":{"slug":"hardware","name":"Hardware & Compute","url":"https://digestai.news/category/hardware"},"entities":{"companies":["AMD","Gufo"],"models":["Qwen3.6-35B-A3B-Q6dense","Qwen3.8-27B-UD-Q8KXL","Qwen3.8-27B-DFlash2-Q4KM"],"people":[]},"firstPublishedAt":"2026-10-03T06:13:45Z","updatedAt":"2026-10-03T06:13:45Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"github.com","title":"gufo-Qwen3.6-35B-A3B-Q6dense - 3095tok/s prefill; 190 tok/s decode on Strix Halo","url":"https://github.com/NinjaPear/gufo-Qwen3.6-35B-A3B-Q6dense","publishedAt":"2026-10-03T06:13:45Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[{"site":"Reddit","url":"https://www.reddit.com/r/LocalLLaMA/comments/1wwfw06/gufoqwen3635ba3bq6dense_3095toks_prefill_190_toks/","points":null}],"thread":{"title":"Open-Source AI Surge Local Models Faster Inference","url":"https://digestai.news/thread/kdnuggets-lists-seven-open-source-chatgpt-alternatives-that-run-locally","storyCount":5},"cite":{"text":"Digest AI, \"gufo-Qwen3.6-35B-A3B-Q6dense - 3095tok/s prefill; 190 tok/s decode on Strix Halo\", 3 October 2026, https://digestai.news/story/gufo-qwen3-6-35b-a3b-q6dense-3095tok-s-prefill-190-tok-s-decode-on-str","publisher":"Digest AI","title":"gufo-Qwen3.6-35B-A3B-Q6dense - 3095tok/s prefill; 190 tok/s decode on Strix Halo","datePublished":"2026-10-03T06:13:45Z","url":"https://digestai.news/story/gufo-qwen3-6-35b-a3b-q6dense-3095tok-s-prefill-190-tok-s-decode-on-str"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}