{"version":1,"type":"story","url":"https://digestai.news/story/ggml-org-halves-indexer-score-memory-in-llama-cpp-pull-request","json":"https://digestai.news/story/ggml-org-halves-indexer-score-memory-in-llama-cpp-pull-request.json","markdown":"https://digestai.news/story/ggml-org-halves-indexer-score-memory-in-llama-cpp-pull-request.md","slug":"ggml-org-halves-indexer-score-memory-in-llama-cpp-pull-request","headline":"Ggml-org halves indexer score memory in llama.cpp Pull Request","summary":"The ggml‑org team submitted a pull request to the llama.cpp repository that changes how the indexer scores are stored during inference. By giving each attention head its own product and summing the results in‑place, the code eliminates the need to keep two large f32 tensors live at once, effectively halving the memory required for the indexer scores.\n\nThe modification uses plain `ggml_add` and `ggml_relu` operations, which the graph allocator already runs in place when there are no other consumers. Compute speed and the compute buffer remain unchanged, and the kernel still supports any head count, with 64 heads running as before. The new approach reads the pooled keys once and avoids per‑head score materialization, reducing memory traffic while keeping the same throughput.\n\nThe change is relevant for models that rely on long‑context inference, such as those built on the ggml tensor library, because it allows larger contexts or batch sizes without increasing GPU or CPU memory usage.","keyPoints":["Memory for indexer scores is halved by storing each head’s product separately and summing in place","Compute speed and buffer usage stay unchanged; 64‑head configurations run as before","The kernel reads pooled keys once, avoiding per‑head score materialization"],"whyItMatters":"Cutting indexer memory in half lets developers run longer‑context models or larger batches on the same hardware, expanding practical AI deployment options.","category":{"slug":"models","name":"Generative AI & Models","url":"https://digestai.news/category/models"},"entities":{"companies":["ggml-org"],"models":[],"people":["CISC","am17an"]},"firstPublishedAt":"2026-10-03T06:18:34Z","updatedAt":"2026-10-03T06:18:34Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"github.com","title":"qwen4exp : halve the indexer score memory by ServeurpersoCom · Pull Request #29825 · ggml-org/llama.cpp","url":"https://github.com/ggml-org/llama.cpp/pull/29825","publishedAt":"2026-10-03T06:18:34Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[{"site":"Reddit","url":"https://www.reddit.com/r/LocalLLaMA/comments/1wwfyv6/qwen4exp_halve_the_indexer_score_memory_by/","points":null}],"thread":{"title":"Llama.cpp Optimizes Qwen4exp Vulkan Inference Performance","url":"https://digestai.news/thread/ggml-org-merges-vulkan-kernel-fusion-for-qwen4exp-s-scalesigmoidscale-chain","storyCount":2},"cite":{"text":"Digest AI, \"Ggml-org halves indexer score memory in llama.cpp Pull Request\", 3 October 2026, https://digestai.news/story/ggml-org-halves-indexer-score-memory-in-llama-cpp-pull-request","publisher":"Digest AI","title":"Ggml-org halves indexer score memory in llama.cpp Pull Request","datePublished":"2026-10-03T06:18:34Z","url":"https://digestai.news/story/ggml-org-halves-indexer-score-memory-in-llama-cpp-pull-request"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}