{"version":1,"type":"story","url":"https://digestai.news/story/github-giveen-ninfer-ext-ninfer-ext-he-built-the-product-we-are-buildi","json":"https://digestai.news/story/github-giveen-ninfer-ext-ninfer-ext-he-built-the-product-we-are-buildi.json","markdown":"https://digestai.news/story/github-giveen-ninfer-ext-ninfer-ext-he-built-the-product-we-are-buildi.md","slug":"github-giveen-ninfer-ext-ninfer-ext-he-built-the-product-we-are-buildi","headline":"GitHub - giveen/ninfer-ext: ninfer-ext — He built the product, we are building the weapon.","summary":"Developer giveen released ninfer-ext, an open-source C++23 and CUDA inference fork based on Neroued's ninfer engine.\n\nThe project fits the large model into consumer hardware by keeping routed experts in pinned host RAM—requiring around 128 GB of system memory and 127 GB of disk space—while streaming active layers over PCIe into a device expert cache. The fork also improves speculative decoding performance, reporting 13% to 48% speedups using DFlash2 on Qwen3.8-27B groupwise-int across one to four concurrent requests compared to stock ninfer.\n\nAdditional updates target agent workflows, including out-of-memory recovery, context-cache salvage for aborted requests, and OpenAI- and Anthropic-compatible local HTTP endpoints. The developer notes that the tool only supports the RTX 5090 architecture and ties upstream performance on certain models like Qwen3.6-35B-A3B.","keyPoints":["Requires 128 GB host RAM and 127 GB disk to offload routed experts over PCIe","Delivers 13% to 48% faster speculative decoding with DFlash2 on Qwen3.8-27B groupwise-int"],"whyItMatters":null,"category":{"slug":"models","name":"Generative AI & Models","url":"https://digestai.news/category/models"},"entities":{"companies":["NVIDIA"],"models":["Qwen3.8-Flash-Next","Qwen3.8-27B","Qwen3.6-35B-A3B","Qwen3.6-27B"],"people":["Neroued"]},"firstPublishedAt":"2026-09-26T19:27:40Z","updatedAt":"2026-09-26T19:27:40Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"github.com","title":"GitHub - giveen/ninfer-ext: ninfer-ext: He built the product, we are building the weapon.","url":"https://github.com/giveen/ninfer-ext","publishedAt":"2026-09-26T19:27:40Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[{"site":"Reddit","url":"https://www.reddit.com/r/LocalLLaMA/comments/1wqz9cr/github_giveenninferext_ninferext_he_built_the/","points":null}],"thread":null,"cite":{"text":"Digest AI, \"GitHub - giveen/ninfer-ext: ninfer-ext — He built the product, we are building the weapon.\", 26 September 2026, https://digestai.news/story/github-giveen-ninfer-ext-ninfer-ext-he-built-the-product-we-are-buildi","publisher":"Digest AI","title":"GitHub - giveen/ninfer-ext: ninfer-ext — He built the product, we are building the weapon.","datePublished":"2026-09-26T19:27:40Z","url":"https://digestai.news/story/github-giveen-ninfer-ext-ninfer-ext-he-built-the-product-we-are-buildi"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}