{"version":1,"type":"story","url":"https://digestai.news/story/llama-cpp-adds-mtp-decoding-for-qwen4exp-and-glm-5-3-flash-hybrid-mode","json":"https://digestai.news/story/llama-cpp-adds-mtp-decoding-for-qwen4exp-and-glm-5-3-flash-hybrid-mode.json","markdown":"https://digestai.news/story/llama-cpp-adds-mtp-decoding-for-qwen4exp-and-glm-5-3-flash-hybrid-mode.md","slug":"llama-cpp-adds-mtp-decoding-for-qwen4exp-and-glm-5-3-flash-hybrid-mode","headline":"Llama.cpp adds MTP decoding for Qwen4Exp and GLM-5.3-Flash hybrid model","summary":"llama.cpp version 0.6.0 introduces **MTP (Mixture of Token Predictors) speculative decoding** for the **Qwen4Exp** model, claiming a **~1.5x speedup** on DGX Spark hardware. The update also adds support for the **GLM-5.3-Flash (GLM5-Next)**, a **320B hybrid text+vision model**, and the **Clef decision model**, which handles both text and vision inputs.\n\nKey improvements include a new **llamabatchext API** for mixed token/embedding inputs, a **Metal tensor-API flash attention kernel** for F16 KV, and **sparse flash attention** for quantized K/V on Vulkan. The project also overhauled its Web UI with a **Hugging Face Hub data layer** and model download pipeline. Under the hood, **ggml** was updated to **v0.26.0**, adding optimizations for CUDA, SYCL, and Vulkan backends, including **BF16 ops** and **MMVQ shared-expert fusion**.","keyPoints":["MTP speculative decoding for Qwen4Exp delivers ~1.5x faster decoding on DGX Spark hardware","GLM-5.3-Flash (320B hybrid text+vision) and Clef decision model now supported","New llamabatchext API, Metal flash attention kernel, and sparse flash attention for Vulkan"],"whyItMatters":"Faster decoding and support for large hybrid models expand llama.cpp’s utility for developers running inference on edge or cloud GPUs. The MTP optimization could reduce latency for real-time applications, while the GLM-5.3-Flash model adds vision capabilities to the toolkit.","category":{"slug":"research","name":"Research","url":"https://digestai.news/category/research"},"entities":{"companies":["llama.cpp"],"models":["Qwen4Exp","GLM-5.3-Flash (GLM5-Next)","Clef","Ling 3.0 VL","Nimble","LFM2.5-Encoder-230M/350M"],"people":[]},"firstPublishedAt":"2026-10-05T18:58:01Z","updatedAt":"2026-10-05T18:58:01Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"github.com","title":"llama.cpp v0.6.0 released with MTP speculative decoding for Qwen4Exp and lots more","url":"https://github.com/ggml-org/llama.cpp/releases/tag/v0.6.0","publishedAt":"2026-10-05T18:58:01Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[{"site":"Reddit","url":"https://www.reddit.com/r/LocalLLaMA/comments/1wyh03u/llamacpp_v060_released_with_mtp_speculative/","points":null}],"thread":{"title":"Optimizing Qwen Models for Local Hardware","url":"https://digestai.news/thread/llama-cpp-expert-pool-fork-for-qwen-3-8-flash-next-iq4-16gb-vram-tested-on-mi50","storyCount":3},"cite":{"text":"Digest AI, \"Llama.cpp adds MTP decoding for Qwen4Exp and GLM-5.3-Flash hybrid model\", 5 October 2026, https://digestai.news/story/llama-cpp-adds-mtp-decoding-for-qwen4exp-and-glm-5-3-flash-hybrid-mode","publisher":"Digest AI","title":"Llama.cpp adds MTP decoding for Qwen4Exp and GLM-5.3-Flash hybrid model","datePublished":"2026-10-05T18:58:01Z","url":"https://digestai.news/story/llama-cpp-adds-mtp-decoding-for-qwen4exp-and-glm-5-3-flash-hybrid-mode"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}