Llama.cpp expert-pool fork for Qwen 3.8 Flash next IQ4 + 16Gb VRAM tested on MI50 gfx906 with parameters
Provides a usable expert‑cache for MoE inference on consumer‑grade AMD GPUs, potentially lowering PCIe traffic and latency for large‑context workloads, though benefits vary by use case.
-
Qwen4Exp adds MTP support in llama.cpp pull request
A pull request merged into the llama.cpp repository on October 1, 2026, adds support for MTP (likely referring to Mixture of Token Predictions or a similar technique) to the Qwen4Exp model. The…
1 source primary source -
Llama.cpp expert-pool fork for Qwen 3.8 Flash next IQ4 + 16Gb VRAM tested on MI50 gfx906 with parameters
An unofficial fork of llama.cpp introduces a persistent expert pool for mixture‑of‑experts (MoE) models that are offloaded with the --n-cpu-moe flag. Upstream llama.cpp lacks an expert cache, and…
1 source primary source