DigestAI news desk

Cut through the AI noise.

Developing story · 2 episodes · 18 Sept to 1 Oct

Llama.cpp expert-pool fork for Qwen 3.8 Flash next IQ4 + 16Gb VRAM tested on MI50 gfx906 with parameters

Provides a usable expert‑cache for MoE inference on consumer‑grade AMD GPUs, potentially lowering PCIe traffic and latency for large‑context workloads, though benefits vary by use case.

  1. Qwen4Exp adds MTP support in llama.cpp pull request

    A pull request merged into the llama.cpp repository on October 1, 2026, adds support for MTP (likely referring to Mixture of Token Predictions or a similar technique) to the Qwen4Exp model. The…

    1 source primary source
  2. Llama.cpp expert-pool fork for Qwen 3.8 Flash next IQ4 + 16Gb VRAM tested on MI50 gfx906 with parameters

    An unofficial fork of llama.cpp introduces a persistent expert pool for mixture‑of‑experts (MoE) models that are offloaded with the --n-cpu-moe flag. Upstream llama.cpp lacks an expert cache, and…

    1 source primary source
Who and what
AtomicChatggml-orgmemoriarumixa3607Qwen3.8-Flash-Next-GGUF, AD-3.84bpw-IQ4XS-M64Qwen4expbyzhangheweiggerganovCISCappro