Nvidia paper shows AI models lose 62.8% accuracy from 4K to 128K tokens
Nvidia researchers released a paper on September 30, 2026 that measured how seven open‑weight language models perform as input length grows from 4,000 to 128,000 tokens. Using a new benchmark called Long‑Transduction, they tested simple tasks such as arithmetic, UUID sorting, variable lookup and table transformation, processing 1,440 documents per model with greedy sampling and exact‑match…
Key points
- Average accuracy across seven models dropped 62.8% when context grew from 4K to 128K tokens.
- DeepSeek fell from 0.909 to 0.554; Nvidia's Nemotron Super fell from 0.711 to 0.138.
- Missing identifiers caused up to a 64.3% performance drop; input format changes caused 36.5% decline.
The study also identified three other factors that hurt performance: changing input formats caused a 36.5% average decline, removing stable identifiers led to up to a 64.3% drop, and increasing local step complexity produced an average 39.9% decline. Nvidia argues that short‑task benchmarks can be misleading and recommends breaking long jobs into smaller, well‑identified subtasks. The findings suggest that pipeline design may be as important as model choice for reliable long‑horizon AI agents.
Nvidia paper finds AI models lose accuracy on long tasks, with steep drops at scale
cryptobriefing.com · 2 October 2026
Loading the full article…
This text was published by cryptobriefing.com and written by Diego Almada Lopez. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- What are world models and how do they predict environments? · 1 src
- Researchers introduce Context language models that manage their own Context · 4 src
- Study finds frontier AI models outperform junior accountants on specific tasks · 2 src
- Qwen4Exp adds MTP support in llama.cpp pull request · 2 src
- Opinion: AlphaGo’s reasoning shows today’s LLMs lack true reasoning · 1 src
Comments
via GitHub Discussions