Ggml-org halves indexer score memory in llama.cpp Pull Request
The ggml‑org team submitted a pull request to the llama.cpp repository that changes how the indexer scores are stored during inference. By giving each attention head its own product and summing the results in‑place, the code eliminates the need to keep two large f32 tensors live at once, effectively halving the memory required for the indexer scores.
Key points
- Memory for indexer scores is halved by storing each head’s product separately and summing in place
- Compute speed and buffer usage stay unchanged; 64‑head configurations run as before
- The kernel reads pooled keys once, avoiding per‑head score materialization
The modification uses plain ggml_add and ggml_relu operations, which the graph allocator already runs in place when there are no other consumers. Compute speed and the compute buffer remain unchanged, and the kernel still supports any head count, with 64 heads running as before. The new approach reads the pooled keys once and avoids per‑head score materialization, reducing memory traffic while keeping the same throughput.
The change is relevant for models that rely on long‑context inference, such as those built on the ggml tensor library, because it allows larger contexts or batch sizes without increasing GPU or CPU memory usage.
The story so far
2 episodes →- Ggml-org halves indexer score memory in llama.cpp Pull Requestthis story
qwen4exp : halve the indexer score memory by ServeurpersoCom · Pull Request #29825 · ggml-org/llama.cpp
github.com · 3 October 2026Loading the full article…
This text was published by github.com. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
Coverage and discussion
1source- Reddit discussionreddit.com
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Generative AI & Models
All →- OpenAI shelves GPT-6.1 Astra after safety and alignment concerns · 3 src
- Anthropic releases Opus 5.5; OpenAI launches Sol and Luna · 1 src
- OpenAI test agents breach Hugging Face after forming 70,000-message coordination board · 1 src
- TypeSafe's Jev decision model costs $0.042 per million tokens, matches Sonnet 5 accuracy · 1 src
- Google releases Gemini 4 Argon to trusted cyber defenders first · 19 src
Comments
via GitHub Discussions