{"version":1,"type":"story","url":"https://digestai.news/story/inco-ai-releases-splash-an-opensource-inference-engine-for-apple-silic","json":"https://digestai.news/story/inco-ai-releases-splash-an-opensource-inference-engine-for-apple-silic.json","markdown":"https://digestai.news/story/inco-ai-releases-splash-an-opensource-inference-engine-for-apple-silic.md","slug":"inco-ai-releases-splash-an-opensource-inference-engine-for-apple-silic","headline":"Inco AI releases Splash, an open‑source inference engine for Apple silicon","summary":"Inco AI announced Splash, a local inference engine built around a specific model and released as open‑source under Apache‑2.0. It runs on Apple silicon Macs with an M3 or newer processor, macOS 26.4 or later, at least 36 GB of unified memory (48 GB recommended), and requires Homebrew. The engine currently supports two models, Qwen3.8-27B and Qwen3.6-35B-A3B, and integrates with LM Studio Bionic for day‑zero use.\n\nBenchmarking on a 48 GB M5 Pro GPU shows Splash decoding up to 2× faster than the next‑fastest engine on Qwen3.8-27B and maintaining a lead across context lengths up to 32K tokens. With four parallel subagents the speedup reaches almost 4×. Prefill throughput reaches about 2,000 input tokens / s on the 35B model and 360 / s on the 27B model, cutting first‑token latency from 29 s to 17 s (35B) and from 317 s to 96 s (27B). Cache‑reuse latency drops to 123 ms (35B) and 282 ms (27B), and concurrency tests completed all 16 simultaneous 32K‑token requests, whereas a general‑purpose engine handled only nine.","keyPoints":["Splash runs on M3 or newer Macs with macOS 26.4+, needs at least 36 GB unified memory (48 GB recommended).","On Qwen3.8-27B, Splash decodes 2× faster than the next‑fastest engine and up to almost 4× faster with four parallel subagents.","The engine is open‑source under Apache‑2.0 at github.com/incoai/splash and currently supports two models."],"whyItMatters":"Local, model‑specific inference can halve latency and boost throughput on consumer Macs, enabling faster AI agents and apps without cloud costs.","category":{"slug":"agents","name":"Agents & Tools","url":"https://digestai.news/category/agents"},"entities":{"companies":["Inco AI","Apple","LM Studio","OpenAI","Anthropic","NVIDIA"],"models":["Qwen3.8-27B","Qwen3.6-35B-A3B"],"people":[]},"firstPublishedAt":"2026-09-17T00:00:00Z","updatedAt":"2026-09-17T00:00:00Z","sourceCount":1,"hasPrimarySource":false,"sources":[{"outlet":"inco.ai","title":"inco.AI: is it worth to jump from non AI-polluted Sequoia to GoldenGate-AI-slop-fest just to run that engine?","url":"https://inco.ai/blog/splash","publishedAt":"2026-09-17T00:00:00Z","type":"press","primary":false,"lead":true}],"sourceNotes":null,"discussions":[{"site":"Reddit","url":"https://www.reddit.com/r/LocalLLaMA/comments/1wlhqbr/incoai_is_it_worth_to_jump_from_non_aipolluted/","points":null}],"thread":null,"cite":{"text":"Digest AI, \"Inco AI releases Splash, an open‑source inference engine for Apple silicon\", 17 September 2026, https://digestai.news/story/inco-ai-releases-splash-an-opensource-inference-engine-for-apple-silic","publisher":"Digest AI","title":"Inco AI releases Splash, an open‑source inference engine for Apple silicon","datePublished":"2026-09-17T00:00:00Z","url":"https://digestai.news/story/inco-ai-releases-splash-an-opensource-inference-engine-for-apple-silic"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}