# Ggml-org halves indexer score memory in llama.cpp Pull Request

Digest AI · Generative AI & Models · published 2026-10-03T06:18:34Z

Canonical: https://digestai.news/story/ggml-org-halves-indexer-score-memory-in-llama-cpp-pull-request

## Summary

The ggml‑org team submitted a pull request to the llama.cpp repository that changes how the indexer scores are stored during inference. By giving each attention head its own product and summing the results in‑place, the code eliminates the need to keep two large f32 tensors live at once, effectively halving the memory required for the indexer scores.

The modification uses plain `ggml_add` and `ggml_relu` operations, which the graph allocator already runs in place when there are no other consumers. Compute speed and the compute buffer remain unchanged, and the kernel still supports any head count, with 64 heads running as before. The new approach reads the pooled keys once and avoids per‑head score materialization, reducing memory traffic while keeping the same throughput.

The change is relevant for models that rely on long‑context inference, such as those built on the ggml tensor library, because it allows larger contexts or batch sizes without increasing GPU or CPU memory usage.

## Key points

- Memory for indexer scores is halved by storing each head’s product separately and summing in place
- Compute speed and buffer usage stay unchanged; 64‑head configurations run as before
- The kernel reads pooled keys once, avoiding per‑head score materialization

## Why it matters

Cutting indexer memory in half lets developers run longer‑context models or larger batches on the same hardware, expanding practical AI deployment options.

## Sources

1. [qwen4exp : halve the indexer score memory by ServeurpersoCom · Pull Request #29825 · ggml-org/llama.cpp](https://github.com/ggml-org/llama.cpp/pull/29825) (github.com, 2026-10-03, primary source)

Part of the developing story: [Llama.cpp Optimizes Qwen4exp Vulkan Inference Performance](https://digestai.news/thread/ggml-org-merges-vulkan-kernel-fusion-for-qwen4exp-s-scalesigmoidscale-chain) (2 stories)

## Cite

Digest AI, "Ggml-org halves indexer score memory in llama.cpp Pull Request", 3 October 2026, https://digestai.news/story/ggml-org-halves-indexer-score-memory-in-llama-cpp-pull-request

---

Written by Digest AI's editorial model from the linked sources; the sources are the record. Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse
JSON: https://digestai.news/story/ggml-org-halves-indexer-score-memory-in-llama-cpp-pull-request.json
