gufo-Qwen3.6-35B-A3B-Q6dense - 3095tok/s prefill; 190 tok/s decode on Strix Halo
The Gufo project posted benchmark results for its vertical local inference engine built for AMD Strix Halo hardware (Ryzen AI MAX+ 395 with Radeon 8060S, up to 128 GiB unified memory). The engine uses speculative decoding with a DFlash2 draft model and targets rootless Podman containers on Linux x86-64 with ROCm 7.2.3. Gufo provides an OpenAI-compatible API at localhost:8080/v1 and supports…
Key points
- Runs Qwen3.6-35B-A3B-Q6dense with DFlash2 speculative decoding via Podman
- OpenAI-compatible API on localhost:8080/v1; MIT-licensed code, separate model weights
The story so far
5 episodes →- gufo-Qwen3.6-35B-A3B-Q6dense - 3095tok/s prefill; 190 tok/s decode on Strix Halothis story
gufo-Qwen3.6-35B-A3B-Q6dense - 3095tok/s prefill; 190 tok/s decode on Strix Halo
github.com · 3 October 2026Loading the full article…
This text was published by github.com. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
Coverage and discussion
1source- Reddit discussionreddit.com
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Hardware & Compute
All →- SpaceX launches Google AI chips into orbit for Project Suncatcher test · 12 src
- Volantis raises $88M Series A to build photonic AI inference system · 3 src
- Nvidia-backed Nebius buys Inferize to cut AI inference GPU costs · 2 src
- NVIDIA Blackwell GPUs power OpenAI's GPT-6 Astra Ultrafast with up to 8x faster token generation · 4 src
- Meta releases open-source Muse Gadgets SDK and free Muse Home Link for U.S. subscribers · 5 src
Comments
via GitHub Discussions