DigestAI news desk

Cut through the AI noise.

Generative AI & Models2 min read

Ggml-org halves indexer score memory in llama.cpp Pull Request

The ggml‑org team submitted a pull request to the llama.cpp repository that changes how the indexer scores are stored during inference. By giving each attention head its own product and summing the results in‑place, the code eliminates the need to keep two large f32 tensors live at once, effectively halving the memory required for the indexer scores.

1 source primary source

Key points

  • Memory for indexer scores is halved by storing each head’s product separately and summing in place
  • Compute speed and buffer usage stay unchanged; 64‑head configurations run as before
  • The kernel reads pooled keys once, avoiding per‑head score materialization

The modification uses plain ggml_add and ggml_relu operations, which the graph allocator already runs in place when there are no other consumers. Compute speed and the compute buffer remain unchanged, and the kernel still supports any head count, with 64 heads running as before. The new approach reads the pooled keys once and avoids per‑head score materialization, reducing memory traffic while keeping the same throughput.

The change is relevant for models that rely on long‑context inference, such as those built on the ggml tensor library, because it allows larger contexts or batch sizes without increasing GPU or CPU memory usage.

The story so far

2 episodes →
  1. Ggml-org halves indexer score memory in llama.cpp Pull Requestthis story
Full story from github.com · via Reddit AI communities primary sourceOpen source ↗

qwen4exp : halve the indexer score memory by ServeurpersoCom · Pull Request #29825 · ggml-org/llama.cpp

github.com · 3 October 2026

Loading the full article…

This text was published by github.com. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗

Coverage and discussion

1source
Topics · follow one to build your own front page
ggml-orgCISCam17an

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Generative AI & Models

All →

Related stories