DigestAI news desk

Cut through the AI noise.

Hardware & Compute7 min read

[For 780M] Vulkan inference 12.9% faster: MMVQ tuning for AMD 780M/Strix in llama.cpp, Q4_0 ROCmFP4 models arriving one after another

By setting load‑mode to auto, the build now skips memory‑mapped I/O on iGPUs, preventing the model from being loaded twice and halving usable memory.

1 source

New ROCmFP4 quantized models for Strix/Halo are appearing on HuggingFace, including Laguna‑S 2.1 variants and large‑scale MoE models. A known bug with Sliding Window Attention and the preserve‑thinking flag doubles generation time; users should avoid that combination.

The story so far

4 episodes →
  1. [For 780M] Vulkan inference 12.9% faster: MMVQ tuning for AMD 780M/Strix in llama.cpp, Q4_0 ROCmFP4 models arriving one after anotherthis story
Full story from note.com · by 東リ屋 · via Search: LlamaOpen source ↗

[For 780M] Vulkan inference 12.9% faster: MMVQ tuning for AMD 780M/Strix in llama.cpp, Q4_0 ROCmFP4 models arriving one after another

note.com · 26 September 2026

Loading the full article…

This text was published by note.com and written by 東リ屋. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗

Topics · follow one to build your own front page
amdllama.cppq4fp4q4kmq4ksq3kmq40qwen3.5-122b-a10b-mtp

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Hardware & Compute

All →

Related stories