DigestAI news desk

Cut through the AI noise.

Hardware & Compute5 min read

gufo-Qwen3.6-35B-A3B-Q6dense - 3095tok/s prefill; 190 tok/s decode on Strix Halo

The Gufo project posted benchmark results for its vertical local inference engine built for AMD Strix Halo hardware (Ryzen AI MAX+ 395 with Radeon 8060S, up to 128 GiB unified memory). The engine uses speculative decoding with a DFlash2 draft model and targets rootless Podman containers on Linux x86-64 with ROCm 7.2.3. Gufo provides an OpenAI-compatible API at localhost:8080/v1 and supports…

1 source primary source

Key points

  • Runs Qwen3.6-35B-A3B-Q6dense with DFlash2 speculative decoding via Podman
  • OpenAI-compatible API on localhost:8080/v1; MIT-licensed code, separate model weights

The story so far

5 episodes →
  1. gufo-Qwen3.6-35B-A3B-Q6dense - 3095tok/s prefill; 190 tok/s decode on Strix Halothis story
Full story from github.com · via Reddit AI communities primary sourceOpen source ↗

gufo-Qwen3.6-35B-A3B-Q6dense - 3095tok/s prefill; 190 tok/s decode on Strix Halo

github.com · 3 October 2026

Loading the full article…

This text was published by github.com. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗

Coverage and discussion

1source
Topics · follow one to build your own front page
AMDGufoQwen3.6-35B-A3B-Q6denseQwen3.8-27B-UD-Q8KXLQwen3.8-27B-DFlash2-Q4KM

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Hardware & Compute

All →

Related stories