{"version":1,"type":"story","url":"https://digestai.news/story/nvidia-paper-shows-ai-models-lose-62-8-accuracy-from-4k-to-128k-tokens","json":"https://digestai.news/story/nvidia-paper-shows-ai-models-lose-62-8-accuracy-from-4k-to-128k-tokens.json","markdown":"https://digestai.news/story/nvidia-paper-shows-ai-models-lose-62-8-accuracy-from-4k-to-128k-tokens.md","slug":"nvidia-paper-shows-ai-models-lose-62-8-accuracy-from-4k-to-128k-tokens","headline":"Nvidia paper shows AI models lose 62.8% accuracy from 4K to 128K tokens","summary":"Nvidia researchers released a paper on September 30, 2026 that measured how seven open‑weight language models perform as input length grows from 4,000 to 128,000 tokens. Using a new benchmark called Long‑Transduction, they tested simple tasks such as arithmetic, UUID sorting, variable lookup and table transformation, processing 1,440 documents per model with greedy sampling and exact‑match scoring. The average accuracy fell 62.8% across the length increase, with individual models showing sharper drops: DeepSeek fell from 0.909 to 0.554, while Nvidia’s own Nemotron Super dropped from 0.711 to 0.138.\n\nThe study also identified three other factors that hurt performance: changing input formats caused a 36.5% average decline, removing stable identifiers led to up to a 64.3% drop, and increasing local step complexity produced an average 39.9% decline. Nvidia argues that short‑task benchmarks can be misleading and recommends breaking long jobs into smaller, well‑identified subtasks. The findings suggest that pipeline design may be as important as model choice for reliable long‑horizon AI agents.","keyPoints":["Average accuracy across seven models dropped 62.8% when context grew from 4K to 128K tokens.","DeepSeek fell from 0.909 to 0.554; Nvidia's Nemotron Super fell from 0.711 to 0.138.","Missing identifiers caused up to a 64.3% performance drop; input format changes caused 36.5% decline."],"whyItMatters":"The results highlight that scaling token windows does not guarantee reliable agent behavior, urging developers to redesign pipelines and labeling strategies for long‑running AI tasks.","category":{"slug":"research","name":"Research","url":"https://digestai.news/category/research"},"entities":{"companies":["Nvidia","DeepSeek"],"models":["Nemotron Super","DeepSeek"],"people":[]},"firstPublishedAt":"2026-10-02T04:58:00Z","updatedAt":"2026-10-02T04:58:00Z","sourceCount":1,"hasPrimarySource":false,"sources":[{"outlet":"cryptobriefing.com","title":"Nvidia paper finds AI models lose accuracy on long tasks, with steep drops at scale","url":"https://cryptobriefing.com/nvidia-paper-ai-models-long-tasks-accuracy","publishedAt":"2026-10-02T04:58:00Z","type":"press","primary":false,"lead":true}],"sourceNotes":null,"discussions":[],"thread":null,"cite":{"text":"Digest AI, \"Nvidia paper shows AI models lose 62.8% accuracy from 4K to 128K tokens\", 2 October 2026, https://digestai.news/story/nvidia-paper-shows-ai-models-lose-62-8-accuracy-from-4k-to-128k-tokens","publisher":"Digest AI","title":"Nvidia paper shows AI models lose 62.8% accuracy from 4K to 128K tokens","datePublished":"2026-10-02T04:58:00Z","url":"https://digestai.news/story/nvidia-paper-shows-ai-models-lose-62-8-accuracy-from-4k-to-128k-tokens"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}