Anthropic and OpenAI adopt Chinese KV cache tricks, slash cache-read pricing
Western AI companies are now using a set of key‑value cache optimizations that Chinese lab DeepSeek released publicly.
Key points
- Claude Opus 5.5 cuts cache‑read pricing by 60% and GPT‑6.1 Sol by 80% after using the Chinese optimizations
- Smaller cache lowers VRAM needs for long‑context models, improving inference margins for Western AI firms
Anthropic’s Claude Opus 5.5 and OpenAI’s GPT‑6.1 Sol have incorporated these techniques, cutting their cache‑read pricing by 60 % and 80 % respectively versus the prior versions. The lower memory footprint reduces VRAM costs for long‑context inference, making top‑tier models more margin‑positive for the Western providers.
The article notes that Chinese labs have been openly sharing such performance breakthroughs, a shift from earlier “distillation” concerns, and suggests the move may help Western firms stay competitive despite earlier GPU access constraints.
Model pages: Claude Opus 5.5 → · GPT-6.1 Sol → · DeepSeek-V4.1-Flash →
The story so far
4 episodes →- Anthropic and OpenAI adopt Chinese KV cache tricks, slash cache-read pricingthis story
The AI Race Just Got Awkward
insufferable.dev · 30 September 2026
Loading the full article…
This text was published by insufferable.dev and written by insufferable.dev. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
Coverage and discussion
1source- Hacker News discussion · 105 pointsnews.ycombinator.com
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Generative AI & Models
All →- Anthropic launches Claude Sonnet 5.5, claiming 30% faster output and up to 30% lower cost per task than Sonnet 5 · 25 src
- Anthropic launches Claude Opus 5.5 with 40% lower costs and Fable 5.1-level performance · 19 src
- Google releases Gemini 3.8 Flash TTS and Flash-Lite TTS with voice cloning and 100+ languages · 14 src
- OpenAI scraps GPT-6.1 Astra release over safety concerns · 60 src
- Llama.cpp adds GLM-5.3-Flash support in commit 649dcb1 · 1 src
Comments
via GitHub Discussions