DigestAI news desk

Cut through the AI noise.

Research16 min read

Llama.cpp adds MTP decoding for Qwen4Exp and GLM-5.3-Flash hybrid model

llama.cpp version 0.6.0 introduces MTP (Mixture of Token Predictors) speculative decoding for the Qwen4Exp model, claiming a ~1.5x speedup on DGX Spark hardware. The update also adds support for the GLM-5.3-Flash (GLM5-Next), a 320B hybrid text+vision model, and the Clef decision model, which handles both text and vision inputs.

1 source primary source

Key points

  • MTP speculative decoding for Qwen4Exp delivers ~1.5x faster decoding on DGX Spark hardware
  • GLM-5.3-Flash (320B hybrid text+vision) and Clef decision model now supported
  • New llamabatchext API, Metal flash attention kernel, and sparse flash attention for Vulkan

Key improvements include a new llamabatchext API for mixed token/embedding inputs, a Metal tensor-API flash attention kernel for F16 KV, and sparse flash attention for quantized K/V on Vulkan. The project also overhauled its Web UI with a Hugging Face Hub data layer and model download pipeline. Under the hood, ggml was updated to v0.26.0, adding optimizations for CUDA, SYCL, and Vulkan backends, including BF16 ops and MMVQ shared-expert fusion.

Model page: Clef →

The story so far

3 episodes →
  1. Llama.cpp adds MTP decoding for Qwen4Exp and GLM-5.3-Flash hybrid modelthis story
Full story from github.com · via Reddit AI communities primary sourceOpen source ↗

llama.cpp v0.6.0 released with MTP speculative decoding for Qwen4Exp and lots more

github.com · 5 October 2026

Loading the full article…

This text was published by github.com. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗

Coverage and discussion

1source
Topics · follow one to build your own front page
llama.cppQwen4ExpGLM-5.3-Flash (GLM5-Next)ClefLing 3.0 VLNimbleLFM2.5-Encoder-230M/350M

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Research

All →

Related stories