Llama.cpp adds MTP decoding for Qwen4Exp and GLM-5.3-Flash hybrid model
llama.cpp version 0.6.0 introduces MTP (Mixture of Token Predictors) speculative decoding for the Qwen4Exp model, claiming a ~1.5x speedup on DGX Spark hardware. The update also adds support for the GLM-5.3-Flash (GLM5-Next), a 320B hybrid text+vision model, and the Clef decision model, which handles both text and vision inputs.
Key points
- MTP speculative decoding for Qwen4Exp delivers ~1.5x faster decoding on DGX Spark hardware
- GLM-5.3-Flash (320B hybrid text+vision) and Clef decision model now supported
- New llamabatchext API, Metal flash attention kernel, and sparse flash attention for Vulkan
Key improvements include a new llamabatchext API for mixed token/embedding inputs, a Metal tensor-API flash attention kernel for F16 KV, and sparse flash attention for quantized K/V on Vulkan. The project also overhauled its Web UI with a Hugging Face Hub data layer and model download pipeline. Under the hood, ggml was updated to v0.26.0, adding optimizations for CUDA, SYCL, and Vulkan backends, including BF16 ops and MMVQ shared-expert fusion.
Model page: Clef →
The story so far
3 episodes →- Llama.cpp adds MTP decoding for Qwen4Exp and GLM-5.3-Flash hybrid modelthis story
llama.cpp v0.6.0 released with MTP speculative decoding for Qwen4Exp and lots more
github.com · 5 October 2026Loading the full article…
This text was published by github.com. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
Coverage and discussion
1source- Reddit discussionreddit.com
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- Google launches prototype satellite for Project Suncatcher to test AI chips in space · 2 src
- NVIDIA shares streaming robotics pipeline for Cosmos3-DROID dataset · 1 src
- AI agents identify two room-temperature magnetic semiconductor candidates · 1 src
- OpenAI claims to solve Navier–Stokes problem, sparking math community debate · 2 src
- Microsoft Research unveils Quine, a biological world model for life sciences · 1 src
Comments
via GitHub Discussions