[For 780M] Vulkan inference 12.9% faster: MMVQ tuning for AMD 780M/Strix in llama.cpp, Q4_0 ROCmFP4 models arriving one after another
By setting load‑mode to auto, the build now skips memory‑mapped I/O on iGPUs, preventing the model from being loaded twice and halving usable memory.
New ROCmFP4 quantized models for Strix/Halo are appearing on HuggingFace, including Laguna‑S 2.1 variants and large‑scale MoE models. A known bug with Sliding Window Attention and the preserve‑thinking flag doubles generation time; users should avoid that combination.
The story so far
4 episodes →- [For 780M] Vulkan inference 12.9% faster: MMVQ tuning for AMD 780M/Strix in llama.cpp, Q4_0 ROCmFP4 models arriving one after anotherthis story
[For 780M] Vulkan inference 12.9% faster: MMVQ tuning for AMD 780M/Strix in llama.cpp, Q4_0 ROCmFP4 models arriving one after another
note.com · 26 September 2026
Loading the full article…
This text was published by note.com and written by 東リ屋. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Hardware & Compute
All →- Google plans to launch AI chips into space with Project Suncatcher · 1 src
- Broadcom and Marvell both grow from AI demand but differ in valuation · 1 src
- Meta unveils Muse Charm pendant and $1,299 VR Glasses at Connect 2026 · 4 src
- Meta’s Louisiana data center spurs housing boom for 7,500 workers · 1 src
- UK’s £26bn AI supercomputer faces mid-2030s launch delay over power grid limits · 2 src
Comments
via GitHub Discussions