# Llama.cpp adds MTP decoding for Qwen4Exp and GLM-5.3-Flash hybrid model

Digest AI · Research · published 2026-10-05T18:58:01Z

Canonical: https://digestai.news/story/llama-cpp-adds-mtp-decoding-for-qwen4exp-and-glm-5-3-flash-hybrid-mode

## Summary

llama.cpp version 0.6.0 introduces **MTP (Mixture of Token Predictors) speculative decoding** for the **Qwen4Exp** model, claiming a **~1.5x speedup** on DGX Spark hardware. The update also adds support for the **GLM-5.3-Flash (GLM5-Next)**, a **320B hybrid text+vision model**, and the **Clef decision model**, which handles both text and vision inputs.

Key improvements include a new **llamabatchext API** for mixed token/embedding inputs, a **Metal tensor-API flash attention kernel** for F16 KV, and **sparse flash attention** for quantized K/V on Vulkan. The project also overhauled its Web UI with a **Hugging Face Hub data layer** and model download pipeline. Under the hood, **ggml** was updated to **v0.26.0**, adding optimizations for CUDA, SYCL, and Vulkan backends, including **BF16 ops** and **MMVQ shared-expert fusion**.

## Key points

- MTP speculative decoding for Qwen4Exp delivers ~1.5x faster decoding on DGX Spark hardware
- GLM-5.3-Flash (320B hybrid text+vision) and Clef decision model now supported
- New llamabatchext API, Metal flash attention kernel, and sparse flash attention for Vulkan

## Why it matters

Faster decoding and support for large hybrid models expand llama.cpp’s utility for developers running inference on edge or cloud GPUs. The MTP optimization could reduce latency for real-time applications, while the GLM-5.3-Flash model adds vision capabilities to the toolkit.

## Sources

1. [llama.cpp v0.6.0 released with MTP speculative decoding for Qwen4Exp and lots more](https://github.com/ggml-org/llama.cpp/releases/tag/v0.6.0) (github.com, 2026-10-05, primary source)

Part of the developing story: [Optimizing Qwen Models for Local Hardware](https://digestai.news/thread/llama-cpp-expert-pool-fork-for-qwen-3-8-flash-next-iq4-16gb-vram-tested-on-mi50) (3 stories)

## Cite

Digest AI, "Llama.cpp adds MTP decoding for Qwen4Exp and GLM-5.3-Flash hybrid model", 5 October 2026, https://digestai.news/story/llama-cpp-adds-mtp-decoding-for-qwen4exp-and-glm-5-3-flash-hybrid-mode

---

Written by Digest AI's editorial model from the linked sources; the sources are the record. Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse
JSON: https://digestai.news/story/llama-cpp-adds-mtp-decoding-for-qwen4exp-and-glm-5-3-flash-hybrid-mode.json
