Lemonade fixes AMD APU model streaming and drops OpenMOSS ROCm backend
The Lemonade open‑source SDK released version 2026.39.1 and a 2026.40 release candidate on Tuesday. The new 2026.40 RC changes how streaming models are handled on AMD integrated GPUs by allocating memory from the APU GTT pool instead of a fixed vRAM carve‑out. This adjustment lets models such as DeepSeek‑V4‑Flash‑IQ2XXS‑DS4 run on AMD Ryzen AI Max (Strix Halo) where they previously failed.
Key points
- Lemonade 2026.40 RC adds APU GTT‑based streaming support for AMD integrated GPUs.
- OpenMOSS ROCm backend removed because it runs around 40x slower than Vulkan on the same hardware.
- New launch agent for JetBrains’ Junie added; 2026.39.1 introduced configurable vRAM auto‑eviction.
In the same update the OpenMOSS back‑end for ROCm was removed on both Windows and Linux because its performance was reported to be around 40x slower than the Vulkan back‑end on identical hardware. The merge request also added a launch agent for JetBrains’ Junie and carried forward other enhancements like configurable vRAM auto‑eviction introduced in 2026.39.1. The changes are documented on GitHub.
Lemonade fixes AMD APU model streaming, drops OpenMOSS ROCm as ~40x slower than Vulkan
phoronix.com · 23 September 2026Loading the full article…
This text was published by phoronix.com. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
Coverage and discussion
1source- Reddit discussionreddit.com
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Agents & Tools
All →- Amazon blocks Meta's Muse AI agent from shopping on its platform · 24 src
- Microsoft launches Ink Canvas sketchpad with AI image generation on Surface Pro · 1 src
- OpenAI adds GPT-Live voice to ChatGPT app for hands‑free task delegation · 5 src
- YouTube says it will let users build custom feeds using Gemini AI · 4 src
- Meta tests Muse calls that are actually made by humans in a call center · 5 src
Comments
via GitHub Discussions