{"version":1,"type":"story","url":"https://digestai.news/story/llama-cpp-achieves-up-to-42x-faster-prompt-lookup-drafting","json":"https://digestai.news/story/llama-cpp-achieves-up-to-42x-faster-prompt-lookup-drafting.json","markdown":"https://digestai.news/story/llama-cpp-achieves-up-to-42x-faster-prompt-lookup-drafting.md","slug":"llama-cpp-achieves-up-to-42x-faster-prompt-lookup-drafting","headline":"Llama.cpp achieves up to 42x faster prompt lookup drafting","summary":"Benchmarks were run on an Apple M4 Pro (14 cores, 48 GB RAM) using the WikiText‑103 corpus (541 MB) and smaller subsets (25 MB‑200 MB). Median latency per drafted token and memory figures were measured over three runs.\n\nThe work builds on earlier contributions by Daniel Lemire, Martin Ankerl, and a PR from Johannes Gaessler that added a static cache to llama.cpp. All code and results are publicly available in the repository, and the author emphasizes that acceptance rates remain unchanged compared with the original implementation.","keyPoints":["Up to 42× faster prompt‑lookup drafting reported on Apple M4 Pro hardware","Switching outer map to ankerl::unordereddense segmentedmap drives most speed gains"],"whyItMatters":"Faster, leaner prompt‑lookup decoding lets developers run larger language models on consumer‑grade machines, lowering latency and cost for interactive AI applications.","category":{"slug":"research","name":"Research","url":"https://digestai.news/category/research"},"entities":{"companies":["Apple"],"models":[],"people":["Daniel Lemire","Martin Ankerl","Johannes Gaessler"]},"firstPublishedAt":"2026-09-27T00:23:41Z","updatedAt":"2026-09-27T00:23:41Z","sourceCount":1,"hasPrimarySource":false,"sources":[{"outlet":"jadidbourbaki.github.io","title":"42x Faster Prompt Lookup Drafting in llama.cpp","url":"https://jadidbourbaki.github.io/blog/prompt-lookup-llama-cpp","publishedAt":"2026-09-27T00:23:41Z","type":"press","primary":false,"lead":true}],"sourceNotes":null,"discussions":[{"site":"Reddit","url":"https://www.reddit.com/r/LocalLLaMA/comments/1wr5ylm/42x_faster_prompt_lookup_drafting_in_llamacpp/","points":null}],"thread":null,"cite":{"text":"Digest AI, \"Llama.cpp achieves up to 42x faster prompt lookup drafting\", 27 September 2026, https://digestai.news/story/llama-cpp-achieves-up-to-42x-faster-prompt-lookup-drafting","publisher":"Digest AI","title":"Llama.cpp achieves up to 42x faster prompt lookup drafting","datePublished":"2026-09-27T00:23:41Z","url":"https://digestai.news/story/llama-cpp-achieves-up-to-42x-faster-prompt-lookup-drafting"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}