{"version":1,"type":"story","url":"https://digestai.news/story/for-780m-vulkan-inference-12-9-faster-mmvq-tuning-for-amd-780m-strix-i","json":"https://digestai.news/story/for-780m-vulkan-inference-12-9-faster-mmvq-tuning-for-amd-780m-strix-i.json","markdown":"https://digestai.news/story/for-780m-vulkan-inference-12-9-faster-mmvq-tuning-for-amd-780m-strix-i.md","slug":"for-780m-vulkan-inference-12-9-faster-mmvq-tuning-for-amd-780m-strix-i","headline":"[For 780M] Vulkan inference 12.9% faster: MMVQ tuning for AMD 780M/Strix in llama.cpp, Q4_0 ROCmFP4 models arriving one after another","summary":"By setting load‑mode to auto, the build now skips memory‑mapped I/O on iGPUs, preventing the model from being loaded twice and halving usable memory.\n\nNew ROCmFP4 quantized models for Strix/Halo are appearing on HuggingFace, including Laguna‑S 2.1 variants and large‑scale MoE models. A known bug with Sliding Window Attention and the preserve‑thinking flag doubles generation time; users should avoid that combination.","keyPoints":[],"whyItMatters":"Developers can adopt the changes immediately, reducing memory errors and improving throughput for applications that rely on Vulkan-based inference.","category":{"slug":"hardware","name":"Hardware & Compute","url":"https://digestai.news/category/hardware"},"entities":{"companies":["amd","llama.cpp"],"models":["q4fp4","q4km","q4ks","q3km","q40","qwen3.5-122b-a10b-mtp"],"people":[]},"firstPublishedAt":"2026-09-26T04:34:00Z","updatedAt":"2026-09-26T04:34:00Z","sourceCount":1,"hasPrimarySource":false,"sources":[{"outlet":"note.com","title":"[For 780M] Vulkan inference 12.9% faster: MMVQ tuning for AMD 780M/Strix in llama.cpp, Q4_0 ROCmFP4 models arriving one after another","url":"https://note.com/samehadaonsen/n/n43d8cbad039a?hl=en","publishedAt":"2026-09-26T04:34:00Z","type":"press","primary":false,"lead":true}],"sourceNotes":null,"discussions":[],"thread":{"title":"Open-Source AI Surge Local Models Faster Inference","url":"https://digestai.news/thread/kdnuggets-lists-seven-open-source-chatgpt-alternatives-that-run-locally","storyCount":4},"cite":{"text":"Digest AI, \"[For 780M] Vulkan inference 12.9% faster: MMVQ tuning for AMD 780M/Strix in llama.cpp, Q4_0 ROCmFP4 models arriving one after another\", 26 September 2026, https://digestai.news/story/for-780m-vulkan-inference-12-9-faster-mmvq-tuning-for-amd-780m-strix-i","publisher":"Digest AI","title":"[For 780M] Vulkan inference 12.9% faster: MMVQ tuning for AMD 780M/Strix in llama.cpp, Q4_0 ROCmFP4 models arriving one after another","datePublished":"2026-09-26T04:34:00Z","url":"https://digestai.news/story/for-780m-vulkan-inference-12-9-faster-mmvq-tuning-for-amd-780m-strix-i"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}