# gufo-Qwen3.6-35B-A3B-Q6dense - 3095tok/s prefill; 190 tok/s decode on Strix Halo

Digest AI · Hardware & Compute · published 2026-10-03T06:13:45Z · updated 2026-10-03T08:14:27Z

Canonical: https://digestai.news/story/gufo-qwen3-6-35b-a3b-q6dense-3095tok-s-prefill-190-tok-s-decode-on-str

## Summary

The Gufo project posted benchmark results for its vertical local inference engine built for AMD Strix Halo hardware (Ryzen AI MAX+ 395 with Radeon 8060S, up to 128 GiB unified memory). The engine uses speculative decoding with a DFlash2 draft model and targets rootless Podman containers on Linux x86-64 with ROCm 7.2.3. Gufo provides an OpenAI-compatible API at localhost:8080/v1 and supports text, images, tools, and streaming. The project emphasizes quality-preserving optimization, per-model kernels, and concurrent request handling. Windows and dual-Strix RDMA variants exist but are not yet merged. Model weights are downloaded separately from Hugging Face (unsloth and z-lab repos). Build requires GCC 15.3, CMake 3.21+, Ninja, and ROCm development packages. Gufo's code is MIT licensed; model weights retain their publishers' terms.

## Key points

- Runs Qwen3.6-35B-A3B-Q6dense with DFlash2 speculative decoding via Podman
- OpenAI-compatible API on localhost:8080/v1; MIT-licensed code, separate model weights

## Why it matters

Shows high-throughput local LLM inference on consumer AMD APUs with unified memory, enabling private, offline AI workloads without discrete GPUs.

## Sources

1. [gufo-Qwen3.6-35B-A3B-Q6dense - 3095tok/s prefill; 190 tok/s decode on Strix Halo](https://github.com/NinjaPear/gufo-Qwen3.6-35B-A3B-Q6dense) (github.com, 2026-10-03, primary source)
2. [gufo_windows pre-package for Strix Halo users](https://github.com/pixmaate/gufo/releases/tag/windows-2026-10-03) (github.com, 2026-10-03, primary source)

Part of the developing story: [Open-Source AI Surge Local Models Faster Inference](https://digestai.news/thread/kdnuggets-lists-seven-open-source-chatgpt-alternatives-that-run-locally) (5 stories)

## Cite

Digest AI, "gufo-Qwen3.6-35B-A3B-Q6dense - 3095tok/s prefill; 190 tok/s decode on Strix Halo", 3 October 2026, https://digestai.news/story/gufo-qwen3-6-35b-a3b-q6dense-3095tok-s-prefill-190-tok-s-decode-on-str

---

Written by Digest AI's editorial model from the linked sources; the sources are the record. Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse
JSON: https://digestai.news/story/gufo-qwen3-6-35b-a3b-q6dense-3095tok-s-prefill-190-tok-s-decode-on-str.json
